BigID
A discovery platform that finds personal and sensitive data across systems and works out whose data it is.
BigID is a data discovery and privacy platform. It scans systems to locate personal and sensitive information, classifies what it finds, and correlates records so that data belonging to the same individual can be identified across different systems. That correlation is what supports answering requests from individuals about their own data.
A question that sounds simple
An individual writes to an organisation and asks what personal data is held about them.
The obligation to answer is clear in a number of privacy regimes. Meeting it is much harder than it appears, and the difficulty is worth understanding because it explains what this kind of platform is for.
The data is spread across systems. The customer database has a record. The support system has tickets. The marketing platform has a contact. The billing system has invoices. Files somewhere contain a spreadsheet from a project three years ago.
Worse, these systems identify the person differently. One uses a customer number, another an email address, a third a reference generated when they first got in touch.
Answering the question means finding all of it and knowing it is all the same person.
Discovery, then correlation
BigID separates these into two problems, and the second is the distinctive part.
Discovery scans systems and finds where personal and sensitive data is held. It covers databases, file stores, cloud services and applications, which matters because a great deal of personal information sits in documents and spreadsheets rather than in tidy database columns.
Correlation links records across systems, establishing that a row in one and a record in another relate to the same individual.
Discovery alone answers where personal data is. Correlation answers whose it is.
The distinction is easy to gloss over and it decides whether the harder obligations can actually be met. Knowing that a column holds email addresses does not tell you which of them belongs to the person who has written in.
Classification
What is found is labelled by type. Identifiers, financial details, health information, and whatever other categories matter to the organisation.
Classification is what makes findings actionable. A list of thousands of locations is not useful. A statement that these particular systems hold health information, and these hold payment details, supports deciding where attention goes first.
It is also the basis of anything written down afterwards. Records of what personal data an organisation holds are expected under several privacy regimes, and building them from scan findings gives a firmer foundation than assembling them from what teams remember.
What this makes possible
Individual rights requests. When someone asks what is held about them, the correlated view is used to compile it, rather than sending the question to a dozen system owners and hoping.
Records of processing. Documentation of what is held and where can be grounded in findings.
Retiring systems. Before a system is switched off, establishing what personal data it holds informs whether it can be deleted, must be retained or needs migrating.
Prioritising protection. Knowing which stores hold the most sensitive information tells you where controls matter most.
The uncomfortable part
Scanning tends to find personal data in places it should not be.
Extracts saved to shared drives. Test environments populated with production data. Backups of systems that were retired years ago. Spreadsheets attached to old emails and saved into a folder.
This is the normal result rather than a sign of unusual mismanagement, and it is the reason discovery is worth doing. Data nobody knows about cannot be protected, deleted on request or reported accurately.
Organisations should expect findings that require action, and plan for that before scanning rather than being surprised by it.
Who uses it
BigID is used by privacy teams, data protection officers, security teams and data governance functions, particularly in organisations holding personal data about large numbers of individuals and subject to privacy regulation.
Points to consider
BigID is a commercial platform. The official documentation is the reference for supported systems, deployment options and capabilities.
Discovery finds what it is configured and able to recognise. Coverage depends on which systems are connected, and systems left unconnected remain unknown regardless of how thorough the scanning is elsewhere.
Correlation across systems is an inference. It links records based on the evidence available, and reviewing how confident those links are is part of using the results responsibly, particularly when compiling a response to an individual.
Findings create obligations. Once an organisation knows personal data sits somewhere it should not, that knowledge is itself significant. This is an argument for doing the work with a plan for acting on the results, not an argument against doing it.
Getting started
The official documentation covers connecting systems, configuring scans, classification and the correlation capability. Scanning a single well understood system first is a sensible way to calibrate what the results look like before pointing it at an entire estate.
Key features of BigID
Capabilities described in the official documentation.
Scanning across many system types
Databases, file stores, cloud services and applications are examined to locate personal information.
Correlation to individuals
Records in different systems are linked so that data about one person can be identified as such.
Classification by data type
Discovered values are labelled by what they are, which is what policies and reports refer to.
Support for individual rights requests
The correlated view is used to compile what is held about a person when they ask.
Advantages of BigID
Factual advantages that follow from the features above.
The actual locations become known
Scanning finds personal data in places no register recorded, which is where exposure usually sits.
Requests about a person can be answered
Correlation means data about an individual is assembled rather than searched for system by system.
Reporting rests on findings
Statements about what is held are based on scan results rather than on what people remember.
Risk can be prioritised
Knowing what type of data sits where allows attention to go to the most sensitive holdings first.
Common use cases for BigID
Situations the official documentation describes this tool as being used for.
Answering subject access requests
When an individual asks what is held about them, the correlated view is used to compile the answer.
Building a record of processing
Discovery findings inform documentation of what personal data the organisation actually holds.
Locating data before a system is retired
Scanning establishes what personal information a system holds before it is decommissioned.
Finding data in unstructured stores
File shares and document repositories are examined for personal details held outside databases.
Official website
Everything on this page is based on the official documentation for BigID. You can read the source here.
BigID official documentation