Detecting personal data (shadow data) - or how to find personal data that a company is unaware of
Shadow data is personal data outside your inventory, retention and safeguards. How to detect it and why visibility is a prerequisite for GDPR compliance.
Most organisations can identify where the personal data they process is located: in their CRM, HR system or ERP. The problem is that the actual location of the data often does not match its description in the record of processing activities. A copy of a report in an email inbox, a 'quick' export to a spreadsheet, an attachment on a shared drive, a folder belonging to an employee who left two years ago - all of these constitute personal data that the data controller does not formally see. This is shadow data: data outside the inventory, outside retention policies and outside any security measures - because you cannot protect something you do not know exists. For the Data Protection Officer, the legal team and senior management - who today bear personal responsibility for information security - this is not merely an organisational issue, but an undisclosed risk that no one has assessed or recorded in the risk register.
Where does shadow data come from and why is it not visible in the ROPA
Shadow data is personal data that is generated and transferred outside official, managed systems, databases and repositories. A similar phenomenon is shadow AI, i.e. the use of AI tools by employees outside the organisation's official oversight. In practice, employees use such unauthorised AI tools for work purposes and often enter personal data into them, which increases the risk of a personal data breach. The data does not end up there as a result of an attack or deliberate circumvention of procedures - it arises as a natural by-product of day-to-day work. A few mechanisms account for the majority of cases: Drafts and exports - an analyst downloads a list of clients into a spreadsheet to prepare a report. HR sends the payroll list by email to the accounts department. A sales representative saves a contact database locally 'just in case'. A record in the production system has an owner, a purpose and a retention period - its copy in a file has none of these attributes. Email and shared repositories - email inboxes and repositories such as SharePoint or Google Drive, network drives or cloud folders systematically accumulate attachments containing personal data. There is no mechanism to delete them - the file survives there for much longer than the declared retention period, often even after the source record has been 'deleted' from the production system. Orphaned data - files left behind by employees who have left the company, directories of closed projects, and archives of applications that have been decommissioned. Nobody reviews them because nobody feels responsible for them. The common denominator is that the record of processing activities describes processes and systems, whilst shadow data exists in files and across systems. The record of processing activities may be maintained to the highest standard and still fail to cover any of these resources, as they, by definition, arise outside the described process. This is why one of the most common findings of audits is a discrepancy between what an organisation declares and what it actually holds - and the most common source of this discrepancy is precisely unstructured data. The practical implication is that the principles of data retention limitation, accuracy and data minimisation (Article 5(1)(c), (d) and (e) of the GDPR) cannot be applied to a layer that is not visible. The statement "we only process necessary data" is true only for systems covered by the inventory. Everything outside this scope is an area where the principles formally apply but, in practice, do not work. Therefore, detecting shadow data is not merely an IT housekeeping task - it is a prerequisite for ensuring that processing complies with the GDPR, including accountability under Article 5(2) of the GDPR.
How to detect personal data that is not inventoried
Unfortunately, a manual review is not an option, as it is impossible for a team to manually review millions of files on shared drives and in mailboxes; detecting shadow data requires a sophisticated tool architecture. The detection of personal data is based on the systematic scanning of distributed resources and the classification of their content: personal data is recognised on the basis of patterns (regular expressions for PESEL numbers, NIP numbers or bank account numbers in IBAN format), dictionaries and heuristics, whilst contextual analysis improves the effectiveness of detection. Machine learning (ML) and natural language processing (NLP) technologies are used to analyse documents, including in scenarios involving the use of data and AI systems, and the results map the data to specific locations, processes and owners. It is advisable to carry out scanning in two modes. A one-off preventive audit provides a snapshot of the current state of affairs - a starting point for organising assets. Continuous monitoring detects new occurrences and data growth before they reach a scale that is difficult to manage, as shadow data is not a one-off phenomenon: it is generated daily, with every export and attachment. The greatest areas of risk remain employees' email inboxes and devices - treated as private workspaces, rarely subject to routine audits, and in practice full of documents containing customer data. Scans of such locations regularly reveal hundreds, or even thousands, of files containing names and account numbers, the existence of which the IT department was completely unaware of. Behavioural analysis also enables the detection of unauthorised data transfers to SaaS applications and other AI-based tools, including unauthorised channels for document analysis. It is crucial that the result is actionable: not 'personal data is somewhere', but 'these files, in these locations, contain these categories of data, and this person is responsible for them'. Without such granularity and detail, it is impossible to make any decision on how to handle the data.
Excessive processing as a risk - particularly in the event of an incident
From the perspective of an information security and privacy management system, any set of personal data stored without a legitimate need is an asset subject to risk assessment on the same terms as other information. Excess data is not business-neutral either - it is a pure cost factor in business operations: it generates risk without delivering value. However, the risk posed by shadow data extends not only to personal data but also to breaches of business confidentiality and the disclosure of trade secrets. This is most evident in scenarios involving cybersecurity incidents and personal data breaches. The scope of a breach is determined by the scope of the data it covers. In the event of a successful ransomware attack or data exfiltration, the scale of the breach is equal to the volume and types of personal data to which the attacker gained access - including data the organisation was unaware of. Using personal accounts instead of work accounts is particularly dangerous, as it increases the risk of data leaks. For example, every forgotten file containing customer data increases the severity of the breach and the potential consequences under Article 83 of the GDPR, including the amount of the financial penalty, without offering anything in return. Shadow data hinders an effective response to an incident. This is the most serious consequence, which is often overlooked. Article 33 of the GDPR requires a data breach to be reported to the supervisory authority - where there is a risk - within 72 hours, whilst Article 34 requires data subjects to be notified where there is a high risk to their rights and freedoms. Both obligations presuppose that the controller is able to determine the scope and assess the scale of the breach: which categories of data, how many individuals, and in which systems. This cannot be done for data whose existence is unknown. The organisation is then faced with a choice between two bad options: to report the breach with an underestimated scope (risking regulatory penalties and a loss of credibility with the UODO and customers) or to admit that it is unable to determine who has been affected and to what extent (which in itself is an incriminating admission). Added to this are reporting obligations to the CSIRT under the National Cybersecurity System Act, and in the financial sector - the ICT incident reporting regime under the DORA Regulation. The conclusion for an organisation's management is an uncomfortable one: it is impossible to secure, minimise or promptly delete data that one is unaware of. Visibility is a prerequisite for any proportionate action and for exercising due diligence.
From detection to mitigation - what to do with the findings
Detection is necessary but not sufficient. A map of redundant data that is not followed up with action remains the same attack surface - only now it is documented. Every finding requires a decision: whether to delete the file, anonymise it, subject it to a retention policy with a mandatory automatic deletion deadline, or - if the data is still necessary - secure the data and restrict access to it in accordance with the need-to-know principle. The classification of information should cover all data, including personal data and special categories of data. In ISO 27001 (Information Security Management System) terminology, minimisation is a form of risk reduction achieved by removing an asset (in this case, a key asset, i.e. information), rather than merely adding a security measure - less data means less risk to manage in the first place. This is supplemented by specific safeguards set out in Annex A to the standard: information deletion, data masking and prevention of data leaks. Organisations should carry out risk assessments and evaluate the risks associated with the implementation or use of ICT systems in the context of the GDPR and regulatory compliance. Simply declaring a retention period without an automatic mechanism to enforce the deletion or anonymisation of data once that period has elapsed is, in practice, a procedural fiction confirmed by regulatory documents - the data remains anyway and continues to be processed, because those responsible for deleting it fail to remember to do so manually. Shadow data is one of the few areas where legal obligations and the economics of security point in exactly the same direction and provide the same guidelines for action: less data outside one's control means both a lower risk of penalties and a smaller attack surface. Many organisations still consider themselves unprepared in this regard, so formalising policies and implementing systematic measures are not an overreaction here, but rather rational and justified steps towards enhancing information cybersecurity. For management, this is an argument for treating the identification of unstructured data not as a mere tidying-up exercise, but as an item on the risk register and an action in risk management plans - ideally before these risks materialise in the most catastrophic scenarios of incidents and breaches.