First page of Microsoft's 100,000-partner directory, sorted by responsiveness All 6 Microsoft Solutions Partner designations Microsoft Solutions Partner since 2006 1,100+ organizations under management
Home/Services/GDPR Data Discovery Service
Training

GDPR Data Discovery Service — Personal Data Inventory & PII Discovery

GDPR Data Discovery Service is a 2 - 6 weeks engagement from IT Partner for organizations that need to identify and inventory personal data relevant to the GDPR across online and on-premises data sources. The service helps answer how much data exists, where it resides, and which data may be impacted by GDPR, including data that contains PII. SKU: ITPWW190TRNOT. Price: $15,000. Manager: Mike Mackey. This service is not a tool to certify GDPR compliance.

Timeline 2 - 6 weeksService owner Mike MackeyOffice 365microsoft 365

What this engagement is

This service identifies and creates an inventory of personal data relevant to the GDPR. IT Partner discovers and scans data sources within the organization, documents where data resides, and provides insight into how much of the data contains personally identifiable information (PII) or sensitive personal information that might be subject to the GDPR. IT Partner also guides customers to appropriate solutions that can help elevate GDPR maturity and work toward GDPR compliance. The GDPR Discovery Toolkit is NOT a tool to certify compliance; it is the distributed responsibility of the customer and their legal and compliance teams to certify their own GDPR compliance. SKU: ITPWW190TRNOT. Price: $15,000. Duration: 2 - 6 weeks. Manager: Mike Mackey.

Success criteria

01How much data do I have?
02Where is my data?
03Which data is impacted by the GDPR? (For example, if it contains PII - Personal Identifiable Information)

What you receive

Data source inventory. You'll be able to identify and document the data sources within the organization online and on-premises
A catalog of discovered data sources used as input for the next steps, including the exact location and access type for every data source, together with information about the type of data and its owner
Classification and Labeling. This helps in facilitating the discussion with the legal, business, and IT teams
Defining a data classification taxonomy and label schema based on legal and business requirements
A finalized classification scheme that defines policies, labels, and conditions
Documentation of the classification scheme for further processing in consecutive steps
Reporting on data discovery

How the work unfolds

Kick-off meeting

During this meeting, the team members will be introduced and the team will be briefed on the upcoming activities, proposed timelines, and expected outcome. IT Partner will explain the technical details of the scanning process, how it works, and what it scans for. Important topics to discuss include, but are not limited to, the scope of the discovery process, the required prerequisites, and the expected outcome.

Step 1 -- Identify and assess

The first activity is to identify all possible data sources across the organization. The data sources that are discovered will be documented in a catalog that will be used as input for the next steps. For every data source, the exact location and access type will be recorded, together with information about the type of data and its owner.

Step 2 -- Plan for automated classification and labeling

After identifying all possible data sources, the next step towards GDPR compliance is to assess whether the GDPR will apply to the organization, and if so, to what extent. Microsoft offers two solutions for automated data discovery and classification: Microsoft Information Protection (MIP) and Microsoft Advanced Data Governance (ADG). For both solutions, the automated classification of documents will rely on the configuration of policies, labels, and conditions. As a result, IT Partner will translate legal and business requirements into a finalized classification scheme that defines policies, labels, and conditions, and document the classification scheme for further processing in consecutive steps.

Step 3 -- Select the right solution

Microsoft currently offers two solutions for data governance: Microsoft Information Protection and Microsoft Advanced Data Governance. Although both support data classification, automated labeling based on sensitive data types, and other data protection and security features, each solution has its own specific deployment scenarios and, at the time of this writing, the two solutions are not fully integrated yet.

-- Confirm discovery path and scan approach

IT Partner and the customer confirm which discovery paths apply to the agreed scope. The online path is used for Microsoft 365 workloads such as Exchange Online mailboxes, SharePoint Online, OneDrive for Business, Microsoft 365 Groups, Teams-connected content, and public folders where permissions and licensing are available. The on-premises path is used for supported file shares, local folders on the scanner server, and supported SharePoint Server sites where Microsoft Information Protection scanner prerequisites are met. Both paths may be included in the same engagement when both online and on-premises repositories are in scope. Non-Microsoft, offline, endpoint, removable media, backup, or application-specific repositories may require manual inventory, export-based review, sampling, or separately scoped tooling if they cannot be scanned directly with the Microsoft capabilities used in this service.

Step 4 (on-premises) -- Implement Microsoft Information Protection

The MIP scanner service will be installed on an on-premises domain joined member server and relies on the cloud-based Microsoft Information Protection service for policy and labeling information. This type of operation requires an integrated deployment in which both the on-premises Active Directory and the Azure Active Directory are synchronized.

Step 5 (on-premises) -- Scan existing data sources

Microsoft Information Protection (MIP) Scanner lets you discover, classify, and protect files on local folders on the Windows Server computer that runs the scanner, network shares through UNC paths that use the Common Internet File System (CIFS) protocol, and sites and libraries for SharePoint Server 2016 and SharePoint Server 2013.

Step 4 (online) -- Prepare for Search and Discovery

Recognition of Personal Identifiable Information (PII) in Microsoft 365 data sources relies on automated data classification through recognition of sensitive data types. The data classification process can be automated so that existing data is scanned and labeled automatically, and new data is classified upon creation. Microsoft 365 Labels and Advanced Data Governance (ADG) will be used for data classification and labeling. Classification can be done automatically or by enabling the user to apply a label to content manually. Once labeled and classified, the data can then be protected by additional means or searched for the presence of specific words or data types.

Step 5 (online) -- Search and Discovery in Microsoft 365

IT Partner will use the Content Search feature in the Microsoft 365 Security & Compliance Center to create the inventory. Content Search allows you to search all content locations in your Microsoft 365 organization, including all mailboxes, inactive mailboxes, the mailboxes for all Microsoft 365 Groups and Microsoft Teams, all SharePoint and OneDrive for Business sites, the sites for all Microsoft 365 groups and Microsoft Teams, and all public folders.

Prerequisites

Access to data locations, including non-supported locations in OneDrive Personal, Databases (SQL Server, SQL online, Access, other), Local storage (workstation, laptop, tablet, mobile phone), Removable devices (USB drives, mobile hard disks), Backup media/3rd-party applications (Non-Microsoft cloud solutions)
Readiness to implement Microsoft Information Protection
For the MIP scanner service: an on-premises domain joined member server
For on-premises Microsoft Information Protection scanner operation: an integrated deployment in which both the on-premises Active Directory and the Azure Active Directory are synchronized
-- Confirm in-scope repositories, business units, countries/regions, data owners, and any exclusions before discovery begins.
-- Provide customer project contacts for IT, security, compliance/privacy, legal, and business data ownership decisions.
-- Provide Microsoft 365 administrative, compliance, security, or eDiscovery permissions sufficient to configure labels, run Content Search, review discovery output, and export or report results as agreed.
-- Confirm that the Microsoft 365 tenant licensing supports the Microsoft Information Protection, labels, content search, advanced data governance, and compliance capabilities selected for the engagement.
-- Provide least-privilege service accounts or delegated access for in-scope file shares, SharePoint Server locations, databases, application exports, and other repositories that are to be inventoried or scanned.
-- Ensure network connectivity, firewall rules, DNS resolution, and file share permissions allow the scanner server or discovery workstation to reach the approved on-premises data locations.
-- Confirm change-control approval, maintenance windows if required, scan throttling preferences, and any restrictions for production repositories.
-- Provide an initial repository list, available data maps, retention schedules, information classification requirements, and examples of GDPR-relevant personal data to validate search and classification logic.
-- Confirm how discovery results may be stored, who may access them, and whether exports must be encrypted or retained in a specific customer-controlled location.

Who does what

IT Partner

  • Identify what personal data you have and where it resides
  • Report how personal data is used and accessed
  • Recommend how to establish security controls to prevent, detect, and respond to vulnerabilities and data breaches
  • Advise how to keep required documentation and manage data requests and breach notifications

Your team

  • Provide access to data locations (non-supported locations in OneDrive Personal), Databases (SQL Server, SQL online, Access, other), Local storage (workstation, laptop, tablet, mobile phone), Removable devices (USB drives, mobile hard disks), Backup media/3rd-party applications (Non-Microsoft cloud solutions)
  • Resource and activity planning
  • Establish and confirm timelines
  • Readiness to implement Microsoft Information Protection

What's not included

GDPR Discovery Toolkit is NOT a tool to certify compliance.
It is the distributed responsibility of the customer and their legal and compliance teams to certify their own GDPR compliance.
-- Formal legal advice, privacy counsel, Data Protection Officer services, or a legal opinion on GDPR compliance are not included.
-- Full remediation, data cleanup, migration, deletion, archival, retention implementation, or permission restructuring is not included unless separately scoped.
-- Enterprise-wide Microsoft Purview Information Protection rollout, production label enforcement, encryption policy deployment, or end-user adoption program beyond the agreed discovery scope is not included unless separately scoped.
-- Custom connectors, custom application development, third-party discovery tool licensing, database schema remediation, or integration with non-Microsoft repositories is not included unless separately scoped.
-- Processing of data subject access requests, breach response operations, managed compliance monitoring, or ongoing compliance operations are not included.
-- Microsoft licensing, Azure consumption, third-party software, hardware, backup media handling, travel, and after-hours work are not included unless explicitly stated in the statement of work.

Limitations & technical notes

!Microsoft Information Protection and Microsoft Advanced Data Governance each have their own specific deployment scenarios and, at the time of this writing, the two solutions are not fully integrated yet.
!The MIP scanner service relies on the cloud-based Microsoft Information Protection service for policy and labeling information.
!On-premises MIP scanner operation requires an integrated deployment in which both the on-premises Active Directory and the Azure Active Directory are synchronized.
!Microsoft Information Protection (MIP) Scanner lets you discover, classify, and protect files on the following data stores: local folders on the Windows Server computer that runs the scanner; network shares through UNC paths that use the Common Internet File System (CIFS) protocol; and sites and libraries for SharePoint Server 2016 and SharePoint Server 2013.
!Recognition of Personal Identifiable Information (PII) in Microsoft 365 data sources relies on automated data classification through recognition of sensitive data types.
!Content Search in the Microsoft 365 Security & Compliance Center allows searching all content locations in the Microsoft 365 organization, including all mailboxes, inactive mailboxes, mailboxes for all Microsoft 365 Groups and Microsoft Teams, all SharePoint and OneDrive for Business sites, sites for all Microsoft 365 groups and Microsoft Teams, and all public folders.
!-- Whether the online path, on-premises path, or both paths are used depends on the agreed scope, available access, licensing, repository type, and technical readiness.
!-- Discovery results are indicators of likely personal data based on configured sensitive information types, search conditions, labels, indexing, and scan coverage; false positives and false negatives are possible and require customer review.
!-- Encrypted, password-protected, corrupted, unsupported, very large, non-indexed, or inaccessible files may not be fully inspected by Microsoft 365 search or scanner capabilities.
!-- Scanning large repositories may require throttling, staged execution, or extended elapsed time to reduce operational impact and align with Microsoft service limits.
!-- Non-Microsoft cloud services, endpoint-local data, removable media, backup media, and line-of-business applications may require customer-provided exports, separate discovery tooling, or manual inventory methods.
!-- The service provides discovery, inventory, reporting, and recommendations; implementation of all recommended controls is a separate activity unless explicitly included in the statement of work.

Frequently asked questions

What is the GDPR Data Discovery Service?

The GDPR Data Discovery Service is a 2–6 week IT Partner engagement that identifies and inventories personal data relevant to GDPR across online and on-premises data sources. It helps organizations answer how much data they have, where it resides, and which data may be impacted by GDPR, including data containing personally identifiable information (PII). The service SKU is ITPWW190TRNOT.

Does the GDPR Data Discovery Service certify that my organization is GDPR compliant?

No. The GDPR Data Discovery Service, including the GDPR Discovery Toolkit, is not a tool to certify GDPR compliance. IT Partner helps identify personal data, document data locations, and guide next steps, but compliance certification remains the responsibility of the customer and their legal and compliance teams.

What deliverables are included in the GDPR Data Discovery Service?

The service includes a data source inventory, a catalog of discovered data sources, classification and labeling guidance, a data classification taxonomy and label schema, a finalized classification scheme, documentation of that scheme, and reporting on data discovery. The catalog documents details such as data source location, access type, type of data, and data owner where discovered during the engagement.

How long does the GDPR Data Discovery Service take?

The GDPR Data Discovery Service typically takes 2–6 weeks. The actual timeline depends on the agreed discovery scope, availability of required access, readiness for Microsoft Information Protection, and coordination between IT Partner and the customer during planning and execution.

How much does the GDPR Data Discovery Service cost?

The stated price for the GDPR Data Discovery Service is $15,000. If the required scope, environment complexity, or prerequisites differ from the standard service description, the customer should confirm pricing and assumptions with IT Partner before starting the engagement.

What data sources can be included in the discovery?

The service is designed to discover and scan data sources across the organization, including online and on-premises locations. For Microsoft 365, Content Search can search mailboxes, inactive mailboxes, Microsoft 365 Groups and Teams mailboxes, SharePoint and OneDrive for Business sites, Microsoft 365 group and Teams sites, and public folders. For on-premises scanning with Microsoft Information Protection, supported locations include local folders on the scanner server, CIFS/UNC network shares, and SharePoint Server 2016 or SharePoint Server 2013 sites and libraries.

Can the service discover personal data stored outside Microsoft 365?

Yes, the service can include on-premises and non-Microsoft locations in the discovery planning, because customer responsibilities include providing access to databases, local storage, removable devices, backup media, third-party applications, and non-Microsoft cloud solutions. However, support and scanning methods vary by location, so exact coverage for non-supported or non-Microsoft sources should be confirmed with IT Partner during scoping and kickoff.

What happens during the kickoff meeting?

During kickoff, IT Partner introduces the project team, reviews the proposed timeline and expected outcomes, and explains the technical details of the scanning process. The meeting also covers discovery scope, required prerequisites, access needs, and what the scanning process will and will not scan for.

What are the main steps in the GDPR Data Discovery Service?

The engagement starts with kickoff, then identifies and assesses possible data sources across the organization. IT Partner then plans for automated classification and labeling, helps select the appropriate Microsoft data governance approach, and proceeds with the relevant on-premises and/or online discovery activities. The online path uses Microsoft 365 search and discovery capabilities, while the on-premises path can use the Microsoft Information Protection scanner where prerequisites are met.

What Microsoft technologies are used in the service?

The service may use Microsoft Information Protection (MIP), Microsoft Advanced Data Governance (ADG), Microsoft 365 labels, and Content Search in the Microsoft 365 Security & Compliance Center. MIP and ADG both support data classification and automated labeling based on sensitive data types, but they have different deployment scenarios and are not fully integrated according to the service description.

What is required for the Microsoft Information Protection scanner?

For the MIP scanner service, the customer needs an on-premises domain-joined member server. On-premises scanner operation also requires an integrated deployment in which on-premises Active Directory and Azure Active Directory are synchronized, because the scanner relies on the cloud-based Microsoft Information Protection service for policy and labeling information.

What prerequisites must the customer provide before or during the engagement?

The customer must provide access to relevant data locations and be ready to implement Microsoft Information Protection where applicable. Required access may include OneDrive Personal non-supported locations, databases such as SQL Server or Access, local storage on endpoints, removable devices, backup media, third-party applications, and non-Microsoft cloud solutions. Additional prerequisites should be confirmed during kickoff because the source service description does not provide a complete checklist beyond the stated items.

What are IT Partner’s responsibilities in the service?

IT Partner is responsible for helping identify what personal data exists and where it resides, reporting how personal data is used and accessed, and recommending ways to establish security controls to prevent, detect, and respond to vulnerabilities and data breaches. IT Partner also advises on keeping required documentation and managing data requests and breach notifications, within the scope of the discovery engagement.

What are the customer’s responsibilities in the service?

The customer is responsible for providing access to data locations, supporting resource and activity planning, establishing and confirming timelines, and being ready to implement Microsoft Information Protection where applicable. The customer’s legal and compliance teams remain responsible for determining and certifying GDPR compliance.

Will the service classify and label data automatically?

The service plans for automated classification and labeling by translating legal and business requirements into policies, labels, and conditions. Microsoft 365 labels, Microsoft Information Protection, and Advanced Data Governance can support automated or user-applied labeling, but exact implementation depends on the selected solution, deployment scenario, and agreed scope.

Does the service include a data classification taxonomy?

Yes. IT Partner helps define a data classification taxonomy and label schema based on legal and business requirements, then documents a finalized classification scheme that defines policies, labels, and conditions. This is intended to support discussion among legal, business, and IT teams and to provide input for later processing steps.

Will the GDPR Data Discovery Service cause downtime or business interruption?

The service description does not state a planned downtime requirement. Because the engagement involves discovery, scanning, access to data locations, and use of Microsoft 365 or MIP capabilities, any expected business impact should be reviewed during kickoff based on the selected data sources, scan approach, and customer environment.

What happens after the GDPR Data Discovery Service is completed?

After completion, the customer receives discovery reporting, an inventory or catalog of discovered data sources, and documentation of the classification scheme. These outputs can be used as input for next steps such as improving data governance, applying security controls, advancing GDPR maturity, and working with legal and compliance teams on GDPR obligations.

Who should be involved from the customer side?

The service is most effective when IT, legal, business, data owners, and compliance stakeholders participate, because the engagement documents data locations and translates legal and business requirements into classification policies, labels, and conditions. The customer should also provide resources who can grant access to data sources and confirm timelines.

Who manages the GDPR Data Discovery Service at IT Partner?

The listed manager for the GDPR Data Discovery Service is Mike Mackey. Customers should confirm engagement logistics, scheduling, and any environment-specific assumptions with IT Partner before project kickoff.

Didn’t find your question?

Ask it here. A real engineer answers by email within one business day — and if it’s a good one, it becomes part of this page so the next person finds it.

Answered by a person, one time, to your inbox. Nothing you type here is published without a human reviewing and anonymizing it first.

Often combined with

$15,000
2 - 6 weeks
Book a meeting