Skip to main content
BRILLIQS

BigID

A discovery platform that finds personal and sensitive data across systems and works out whose data it is.

BigID is a data discovery and privacy platform. It scans systems to locate personal and sensitive information, classifies what it finds, and correlates records so that data belonging to the same individual can be identified across different systems. That correlation is what supports answering requests from individuals about their own data.

A question that sounds simple

An individual writes to an organisation and asks what personal data is held about them.

The obligation to answer is clear in a number of privacy regimes. Meeting it is much harder than it appears, and the difficulty is worth understanding because it explains what this kind of platform is for.

The data is spread across systems. The customer database has a record. The support system has tickets. The marketing platform has a contact. The billing system has invoices. Files somewhere contain a spreadsheet from a project three years ago.

Worse, these systems identify the person differently. One uses a customer number, another an email address, a third a reference generated when they first got in touch.

Answering the question means finding all of it and knowing it is all the same person.

Discovery, then correlation

BigID separates these into two problems, and the second is the distinctive part.

Discovery scans systems and finds where personal and sensitive data is held. It covers databases, file stores, cloud services and applications, which matters because a great deal of personal information sits in documents and spreadsheets rather than in tidy database columns.

Correlation links records across systems, establishing that a row in one and a record in another relate to the same individual.

Discovery alone answers where personal data is. Correlation answers whose it is.

The distinction is easy to gloss over and it decides whether the harder obligations can actually be met. Knowing that a column holds email addresses does not tell you which of them belongs to the person who has written in.

Classification

What is found is labelled by type. Identifiers, financial details, health information, and whatever other categories matter to the organisation.

Classification is what makes findings actionable. A list of thousands of locations is not useful. A statement that these particular systems hold health information, and these hold payment details, supports deciding where attention goes first.

It is also the basis of anything written down afterwards. Records of what personal data an organisation holds are expected under several privacy regimes, and building them from scan findings gives a firmer foundation than assembling them from what teams remember.

What this makes possible

Individual rights requests. When someone asks what is held about them, the correlated view is used to compile it, rather than sending the question to a dozen system owners and hoping.

Records of processing. Documentation of what is held and where can be grounded in findings.

Retiring systems. Before a system is switched off, establishing what personal data it holds informs whether it can be deleted, must be retained or needs migrating.

Prioritising protection. Knowing which stores hold the most sensitive information tells you where controls matter most.

The uncomfortable part

Scanning tends to find personal data in places it should not be.

Extracts saved to shared drives. Test environments populated with production data. Backups of systems that were retired years ago. Spreadsheets attached to old emails and saved into a folder.

This is the normal result rather than a sign of unusual mismanagement, and it is the reason discovery is worth doing. Data nobody knows about cannot be protected, deleted on request or reported accurately.

Organisations should expect findings that require action, and plan for that before scanning rather than being surprised by it.

Who uses it

BigID is used by privacy teams, data protection officers, security teams and data governance functions, particularly in organisations holding personal data about large numbers of individuals and subject to privacy regulation.

Points to consider

BigID is a commercial platform. The official documentation is the reference for supported systems, deployment options and capabilities.

Discovery finds what it is configured and able to recognise. Coverage depends on which systems are connected, and systems left unconnected remain unknown regardless of how thorough the scanning is elsewhere.

Correlation across systems is an inference. It links records based on the evidence available, and reviewing how confident those links are is part of using the results responsibly, particularly when compiling a response to an individual.

Findings create obligations. Once an organisation knows personal data sits somewhere it should not, that knowledge is itself significant. This is an argument for doing the work with a plan for acting on the results, not an argument against doing it.

Getting started

The official documentation covers connecting systems, configuring scans, classification and the correlation capability. Scanning a single well understood system first is a sensible way to calibrate what the results look like before pointing it at an entire estate.

Key features of BigID

Capabilities described in the official documentation.

Scanning across many system types

Databases, file stores, cloud services and applications are examined to locate personal information.

Correlation to individuals

Records in different systems are linked so that data about one person can be identified as such.

Classification by data type

Discovered values are labelled by what they are, which is what policies and reports refer to.

Support for individual rights requests

The correlated view is used to compile what is held about a person when they ask.

Advantages of BigID

Factual advantages that follow from the features above.

The actual locations become known

Scanning finds personal data in places no register recorded, which is where exposure usually sits.

Requests about a person can be answered

Correlation means data about an individual is assembled rather than searched for system by system.

Reporting rests on findings

Statements about what is held are based on scan results rather than on what people remember.

Risk can be prioritised

Knowing what type of data sits where allows attention to go to the most sensitive holdings first.

Common use cases for BigID

Situations the official documentation describes this tool as being used for.

Retail

Answering subject access requests

When an individual asks what is held about them, the correlated view is used to compile the answer.

Financial services

Building a record of processing

Discovery findings inform documentation of what personal data the organisation actually holds.

Telecommunications

Locating data before a system is retired

Scanning establishes what personal information a system holds before it is decommissioned.

Healthcare

Finding data in unstructured stores

File shares and document repositories are examined for personal details held outside databases.

Official website

Everything on this page is based on the official documentation for BigID. You can read the source here.

BigID official documentation

Frequently asked questions about BigID

Answers taken from the official documentation for this tool.

Discovery finds that a column holds personal data. Correlation links records across systems so it can be established that a particular row here and a particular record there relate to the same individual. Answering what is held about one person requires the second, not only the first.

It supports a range of system types including file stores and document repositories, which matters because a great deal of personal information sits in spreadsheets and documents rather than in databases. The official documentation lists the supported systems.

Because a person's data is spread across systems that identify them differently. One uses an account number, another an email address, a third a customer reference. Assembling everything means knowing those refer to the same individual, which is exactly what correlation establishes.

They overlap and differ in emphasis. A catalogue is primarily about helping people find and understand data for analysis. This is oriented towards locating sensitive and personal information and establishing whose it is, which serves privacy obligations rather than discovery for analysis.