Have you ever thought about how far a data record can travel within your organization?

Take a retailer, for example, holding customer information in the database of their order management system. It gets exported into a file format and uploaded to cloud storage before being shared, downloaded and imported into the marketing CMS for a campaign. 

The data has transited several systems and many users. It’s just as sensitive as when it was captured in the order management database, but is now governed by very different sets of access, ownership and retention controls for each of the environments in which it now sits.

Throughout this journey, the most exposed location may not be the system with the largest volume of sensitive data. It may be the downstream system with broader access and limited oversight visibility that presents the greatest risk.

In this article, we’ll explain how to find and prioritize data risk across these complex, distributed and highly dynamic environments.

Distributed data creates uneven risk

The same customer records can present materially different risks depending on whether they are stored in a controlled database, a restricted analytics platform, an open shared folder or an unmanaged spreadsheet. 

Each of these platforms is governed by a different set of security controls, access restrictions, monitoring and policy enforcement. The governance that’s in place at the source application may not be the same for downstream copies. 

When customer records leave an order-management database as a CSV, the database permissions no longer govern the exported file. Saving it to Google Drive places it under a different identity, sharing and ownership model. Uploading it to a marketing platform creates another version governed through that service’s accounts, roles and retention settings.

This exposes data to differing levels of risk because there’s an inconsistency between the control requirements for main repositories and the subsidiary and incidental systems data comes into contact with through its lifecycle.

For this reason, organizations need to know where their sensitive data is actually located, including all the downstream and incidental locations that copies exist.

Investigate beyond expected systems

Modern businesses are dynamic data environments, with digital information transiting hundreds of systems, people and processes daily. 

Routine business processes create secondary copies through database extracts, application reports, API integrations, email attachments, shared drives and user-created spreadsheets. Data can also remain in archives, backups, legacy repositories and SaaS applications after the workflow or project that created it has ended.

These copies are not necessarily unauthorized. A marketing campaign may legitimately require a customer segment. A finance team may need an extract for reconciliation. An analyst may need to transform operational data for reporting

So how can organizations keep track of their sensitive data effectively?

Many businesses start the process of data discovery by mapping a data flow diagram. These diagrams are useful for showing intended data movement, identifying the primary systems processing specific types of data, but fall short when it comes to prioritizing data risk across complex, dynamic environments.

Sensitive data discovery solutions enable the identification of targeted data types across digital environments, uncovering hidden stores of information created as a byproduct of business processes – the accumulated effects of system integrations, operational processes, workarounds and user behavior.

That is the first requirement for meaningful prioritization: 

Determine where sensitive data actually exists across structured and unstructured environments.

Assess exposure, not just sensitivity

Discovery identifies the data and where risk may exist. In isolation, it doesn’t help establish which findings should be addressed first. 

The sensitivity of the data within a location is relatively inconsequential if that environment is specifically built to host it. However, that same data within a forgotten CSV file in a shared folder with cross-team access may be far more significant. 

The priority of action depends on both sensitivity and the prevailing control environment of the data, taking into account:

  • The type and volume of data

  • The location of the environment 

  • Data ownership records

  • Data age and longevity

  • Security posture of the underlying system/infrastructure

With this information, it’s easier to identify those areas of immediate concern. These are locations where sensitive data is present in the context of broad access, weak ownership and/or vulnerable infrastructure. 

Assess exposure based on sensitivity and the environmental context.

Access is the new perimeter

The security of data is heavily dependent on who is authorized to access it. However, this is the greatest area of variability as data moves from one system to another. 

In our retail example, database records may be restricted to the order management application and its administrators. When these records are exported and uploaded to cloud storage, it’s now available to a wider group of users. Finally, once the information is imported into the marketing system, it’s now accessible to the marketing team and their external agencies.

When evaluating the security posture of sensitive data it’s important to consider the access permissions in place for each version of the data record – from the source to its downstream copies. Evaluating how access changes as the data moves may be an early indicator of risk that needs to be managed. 

Prioritize locations where access permissions are more expansive than those of the primary data source.

Retaining classification and context 

Data classification can be an important security measure for data in transit. Classification labelling works with security solutions like data loss prevention, automation and policy enforcement tools to ensure that sensitive data is prevented from entering unauthorized environments. 

However, as data is exported and manipulated between systems, these labels can be lost. This means that copies of the data do not carry the classification labelling required for solutions like DLP to work effectively, allowing the data to traverse systems and environments where it is not authorized.

Discovery solutions like Enterprise Recon integrate with labeling via Purview/MIP to ensure that these unmarked copies can be identified and classified, helping to limit further unauthorized distribution of sensitive data assets. .

Limit further exposure with data classification labeling.

Prioritizing risk with data-level evidence

Ground Labs Enterprise Recon supports this evidence-led process by discovering sensitive data across supported workstations, servers, databases, email platforms, big-data environments and cloud-storage services. Its risk profiles can use data types, match volumes, location information, metadata and supported access criteria to categorize findings.

For supported file systems, Enterprise Recon can also retrieve and analyze access permissions and metadata for locations containing sensitive-data matches. 

These capabilities provide data-level evidence that helps teams identify their highest areas of risk for targeted remediation.

Turn prioritized risks into targeted actions

The value of risk management comes from the ability to act and address exposures effectively. The scale and complexity of an organization’s data estate means that identifying, assessing and prioritizing that response can be a monumental challenge. 

By combining sensitive data discovery with evidence about location, access, ownership, age and security posture – managed through data intelligence platforms like Enterprise Recon – organizations can focus their limited resources where they can have the greatest impact for a defensible, evidence-based approach to risk reduction across cloud, SaaS and on-premises environments.