First page of Microsoft's 100,000-partner directory, sorted by responsiveness Microsoft Solutions Partner — Security, Modern Work, Infrastructure, App Innovation Microsoft partner since 2006 1,100+ organizations under management
Home/Services/GDPR Data Discovery Service
Training

GDPR Data Discovery Service — Personal Data Inventory & PII Discovery

GDPR Data Discovery Service is a 2 - 6 weeks engagement from IT Partner for organizations that need to identify and inventory personal data relevant to the GDPR across online and on-premises data sources. The service helps answer how much data exists, where it resides, and which data may be impacted by GDPR, including data that contains PII. This service is not a tool to certify GDPR compliance.

Timeline 2 - 6 weeksService owner Mike MackeyOffice 365microsoft 365

What this engagement is

This service identifies and creates an inventory of personal data relevant to the GDPR. IT Partner discovers and scans data sources within the organization, documents where data resides, and provides insight into how much of the data contains personally identifiable information (PII) or sensitive personal information that might be subject to the GDPR. IT Partner also guides customers to appropriate solutions that can help elevate GDPR maturity and work toward GDPR compliance. This service is not a compliance certification; certifying GDPR compliance remains the responsibility of the customer and its legal and compliance teams.

Success criteria

01How much data do I have?
02Where is my data?
03Which data is impacted by the GDPR? (For example, if it contains PII - Personal Identifiable Information)

What you receive

Data source inventory. You'll be able to identify and document the data sources within the organization online and on-premises
A catalog of discovered data sources used as input for the next steps, including the exact location and access type for every data source, together with information about the type of data and its owner
Classification and Labeling. This helps in facilitating the discussion with the legal, business, and IT teams
Defining a data classification taxonomy and label schema based on legal and business requirements
A finalized classification scheme that defines policies, labels, and conditions
Documentation of the classification scheme for further processing in consecutive steps
Reporting on data discovery

How the work unfolds

Kick-off meeting

During this meeting, the team members will be introduced and the team will be briefed on the upcoming activities, proposed timelines, and expected outcome. IT Partner will explain the technical details of the scanning process, how it works, and what it scans for. Important topics to discuss include, but are not limited to, the scope of the discovery process, the required prerequisites, and the expected outcome.

Step 1 -- Identify and assess

The first activity is to identify all possible data sources across the organization. The data sources that are discovered will be documented in a catalog that will be used as input for the next steps. For every data source, the exact location and access type will be recorded, together with information about the type of data and its owner.

Step 2 -- Plan for automated classification and labeling

After identifying all possible data sources, the next step towards GDPR compliance is to assess whether the GDPR will apply to the organization, and if so, to what extent. Classification and labeling use Microsoft Purview Information Protection (sensitivity labels and sensitive information types) together with Microsoft Purview data lifecycle and records management capabilities, configured in the Microsoft Purview portal. Automated classification relies on the configuration of policies, labels, and conditions. As a result, IT Partner will translate legal and business requirements into a finalized classification scheme that defines policies, labels, and conditions, and document the classification scheme for further processing in consecutive steps.

Step 3 -- Confirm Purview capabilities and discovery paths

IT Partner and the customer confirm the Microsoft Purview capabilities and the discovery paths that apply to the agreed scope — sensitivity labels, sensitive information types, and the online and on-premises discovery methods described below.

Confirm discovery path and scan approach

IT Partner and the customer confirm which discovery paths apply to the agreed scope. The online path is used for Microsoft 365 workloads such as Exchange Online mailboxes, SharePoint Online, OneDrive for Business, Microsoft 365 Groups, Teams-connected content, and public folders where permissions and licensing are available. The on-premises path is used for supported file shares, local folders on the scanner server, and supported SharePoint Server sites where Microsoft Information Protection scanner prerequisites are met. Both paths may be included in the same engagement when both online and on-premises repositories are in scope. Non-Microsoft, offline, endpoint, removable media, backup, or application-specific repositories may require manual inventory, export-based review, sampling, or separately scoped tooling if they cannot be scanned directly with the Microsoft capabilities used in this service.

Step 4 (on-premises) -- Implement the Microsoft Purview Information Protection scanner

The Microsoft Purview Information Protection scanner is installed on an on-premises domain-joined member server and relies on the cloud-based Purview Information Protection service for policy and labeling information. This type of operation requires an integrated deployment in which both the on-premises Active Directory and Microsoft Entra ID are synchronized.

Step 5 (on-premises) -- Scan existing data sources

The Microsoft Purview Information Protection scanner discovers and classifies files in local folders on the scanner server, network shares over UNC paths, and supported versions of SharePoint Server.

Step 4 (online) -- Prepare for Search and Discovery

Recognition of Personal Identifiable Information (PII) in Microsoft 365 data sources relies on automated data classification through recognition of sensitive data types. The data classification process can be automated so that existing data is scanned and labeled automatically, and new data is classified upon creation. Microsoft Purview sensitivity labels, retention labels, and sensitive information types will be used for data classification and labeling. Classification can be done automatically or by enabling the user to apply a label to content manually. Once labeled and classified, the data can then be protected by additional means or searched for the presence of specific words or data types.

Step 5 (online) -- Search and Discovery in Microsoft 365

IT Partner will use Content Search inside the unified Microsoft Purview eDiscovery experience to create the inventory — the classic Security & Compliance Center search tools are retired, and content searches now run within Purview eDiscovery in the Purview portal. Content Search allows you to search all content locations in your Microsoft 365 organization, including all mailboxes, inactive mailboxes, the mailboxes for all Microsoft 365 Groups and Microsoft Teams, all SharePoint and OneDrive for Business sites, the sites for all Microsoft 365 groups and Microsoft Teams, and all public folders.

Prerequisites

Access to data locations, including non-supported locations in OneDrive Personal, Databases (SQL Server, SQL online, Access, other), Local storage (workstation, laptop, tablet, mobile phone), Removable devices (USB drives, mobile hard disks), Backup media/3rd-party applications (Non-Microsoft cloud solutions)
Readiness to implement Microsoft Information Protection
For the Microsoft Purview Information Protection scanner: an on-premises domain-joined member server
For on-premises Microsoft Information Protection scanner operation: an integrated deployment in which both the on-premises Active Directory and the Microsoft Entra ID are synchronized
Confirm in-scope repositories, business units, countries/regions, data owners, and any exclusions before discovery begins.
Provide customer project contacts for IT, security, compliance/privacy, legal, and business data ownership decisions.
Provide Microsoft 365 administrative, compliance, security, or eDiscovery permissions sufficient to configure labels, run Content Search, review discovery output, and export or report results as agreed.
Confirm that the Microsoft 365 tenant licensing supports the Microsoft Information Protection, labels, content search, advanced data governance, and compliance capabilities selected for the engagement.
Provide least-privilege service accounts or delegated access for in-scope file shares, SharePoint Server locations, databases, application exports, and other repositories that are to be inventoried or scanned.
Ensure network connectivity, firewall rules, DNS resolution, and file share permissions allow the scanner server or discovery workstation to reach the approved on-premises data locations.
Confirm change-control approval, maintenance windows if required, scan throttling preferences, and any restrictions for production repositories.
Provide an initial repository list, available data maps, retention schedules, information classification requirements, and examples of GDPR-relevant personal data to validate search and classification logic.
Confirm how discovery results may be stored, who may access them, and whether exports must be encrypted or retained in a specific customer-controlled location.

Who does what

IT Partner

  • Identify what personal data you have and where it resides
  • Report how personal data is used and accessed
  • Recommend how to establish security controls to prevent, detect, and respond to vulnerabilities and data breaches
  • Advise how to keep required documentation and manage data requests and breach notifications

Your team

  • Provide access to data locations (non-supported locations in OneDrive Personal), Databases (SQL Server, SQL online, Access, other), Local storage (workstation, laptop, tablet, mobile phone), Removable devices (USB drives, mobile hard disks), Backup media/3rd-party applications (Non-Microsoft cloud solutions)
  • Resource and activity planning
  • Establish and confirm timelines
  • Readiness to implement Microsoft Information Protection

What's not included

Compliance certification — this service is not a tool to certify GDPR compliance.
It is the distributed responsibility of the customer and their legal and compliance teams to certify their own GDPR compliance.
Formal legal advice, privacy counsel, Data Protection Officer services, or a legal opinion on GDPR compliance are not included.
Full remediation, data cleanup, migration, deletion, archival, retention implementation, or permission restructuring is not included unless separately scoped.
Enterprise-wide Microsoft Purview Information Protection rollout, production label enforcement, encryption policy deployment, or end-user adoption program beyond the agreed discovery scope is not included unless separately scoped.
Custom connectors, custom application development, third-party discovery tool licensing, database schema remediation, or integration with non-Microsoft repositories is not included unless separately scoped.
Processing of data subject access requests, breach response operations, managed compliance monitoring, or ongoing compliance operations are not included.
Microsoft licensing, Azure consumption, third-party software, hardware, backup media handling, travel, and after-hours work are not included unless explicitly stated in the statement of work.

Limitations & technical notes

!Automated classification depends on the configuration of policies, labels, and conditions in Microsoft Purview; results are only as good as the agreed classification scheme.
!The Purview Information Protection scanner relies on the cloud-based Purview service for policy and labeling information, and on-premises operation requires Active Directory synchronized with Microsoft Entra ID.
!The scanner covers local folders on the scanner server, network shares over UNC paths, and supported versions of SharePoint Server; other repositories are scoped individually.
!This service is not a compliance certification; GDPR compliance certification remains the customer's responsibility.

Frequently asked questions

What is the GDPR Data Discovery Service?

The GDPR Data Discovery Service is a 2–6 week IT Partner engagement that identifies and inventories personal data relevant to GDPR across online and on-premises data sources. It helps organizations answer how much data they have, where it resides, and which data may be impacted by GDPR, including data containing personally identifiable information (PII).

Does the GDPR Data Discovery Service certify that my organization is GDPR compliant?

No. The GDPR Data Discovery Service is not a compliance certification. IT Partner helps identify personal data, document data locations, and guide next steps, but certifying GDPR compliance remains the responsibility of the customer and its legal and compliance teams.

What deliverables are included in the GDPR Data Discovery Service?

The service includes a data source inventory, a catalog of discovered data sources, classification and labeling guidance, a data classification taxonomy and label schema, a finalized classification scheme, documentation of that scheme, and reporting on data discovery. The catalog documents details such as data source location, access type, type of data, and data owner where discovered during the engagement.

How long does the GDPR Data Discovery Service take?

The GDPR Data Discovery Service typically takes 2–6 weeks. The actual timeline depends on the agreed discovery scope, availability of required access, readiness for Microsoft Information Protection, and coordination between IT Partner and the customer during planning and execution.

How much does the GDPR Data Discovery Service cost?

The stated price for the GDPR Data Discovery Service is $15,000. If the required scope, environment complexity, or prerequisites differ from the standard service description, the customer should confirm pricing and assumptions with IT Partner before starting the engagement.

What data sources can be included in the discovery?

The service is designed to discover and scan data sources across the organization, including online and on-premises locations. For Microsoft 365, Content Search in the unified Microsoft Purview eDiscovery experience can search mailboxes, inactive mailboxes, Microsoft 365 Groups and Teams mailboxes, SharePoint and OneDrive for Business sites, Microsoft 365 group and Teams sites, and public folders. For on-premises scanning with the Microsoft Purview Information Protection scanner, supported locations include local folders on the scanner server, CIFS/UNC network shares, and sites and libraries on supported versions of SharePoint Server.

Can the service discover personal data stored outside Microsoft 365?

Yes, the service can include on-premises and non-Microsoft locations in the discovery planning, because customer responsibilities include providing access to databases, local storage, removable devices, backup media, third-party applications, and non-Microsoft cloud solutions. However, support and scanning methods vary by location, so exact coverage for non-supported or non-Microsoft sources should be confirmed with IT Partner during scoping and kickoff.

What are the main steps in the GDPR Data Discovery Service?

The engagement starts with kickoff, then identifies and assesses possible data sources across the organization. IT Partner then plans for automated classification and labeling, helps select the appropriate Microsoft data governance approach, and proceeds with the relevant on-premises and/or online discovery activities. The online path uses Microsoft 365 search and discovery capabilities, while the on-premises path can use the Microsoft Information Protection scanner where prerequisites are met.

What Microsoft technologies are used in the service?

The service uses Microsoft Purview Information Protection (sensitivity labels and sensitive information types), Microsoft Purview data lifecycle management capabilities, the Purview Information Protection scanner for on-premises sources, and Content Search inside the unified Microsoft Purview eDiscovery experience for Microsoft 365 workloads.

What is required for the Microsoft Purview Information Protection scanner?

For the Microsoft Purview Information Protection scanner, the customer needs an on-premises domain-joined member server. On-premises scanner operation also requires an integrated deployment in which on-premises Active Directory and Microsoft Entra ID are synchronized, because the scanner relies on the cloud-based Purview service for policy and labeling information.

Will the GDPR Data Discovery Service cause downtime or business interruption?

No downtime is planned. Discovery and scanning are read-oriented activities; scans of on-premises sources are scheduled with the customer so any load on file servers or SharePoint farms stays within agreed windows.

What happens after the GDPR Data Discovery Service is completed?

After completion, the customer receives discovery reporting, an inventory or catalog of discovered data sources, and documentation of the classification scheme. These outputs can be used as input for next steps such as improving data governance, applying security controls, advancing GDPR maturity, and working with legal and compliance teams on GDPR obligations.

Didn’t find your question?

Ask it here. A real engineer answers by email within one business day — and if it’s a good one, it becomes part of this page so the next person finds it.

Answered by a person, one time, to your inbox. Nothing you type here is published without a human reviewing and anonymizing it first.

Often combined with

$15,000 per project
2 - 6 weeks
Book a meeting