🚀 CloudSEK featured in the 2026 Frost Radar™!
Read more
A data risk assessment is a structured review that locates an organization's sensitive data, identifies the threats and control gaps around it, and rates each risk by likelihood and impact so remediation follows priority.
It covers data in cloud platforms, SaaS applications, on-premises systems, endpoints, backups, and vendor environments.
The assessment answers 4 practical questions for every important data set: where the data lives, who reaches it, what exposes it, and how much harm exposure causes. Those answers turn a vague sense of data exposure into a ranked list of risks with owners and deadlines.
Regulators and auditors increasingly expect that ranked list as documented evidence of due care. HIPAA, PCI DSS, GDPR, SOC 2, and India's DPDP framework all tie security obligations to a documented assessment of risk.
A completed data risk assessment produces a data inventory, a data flow map, scored risks in a risk register, a treatment plan, and a record of accepted residual risk.
To conduct a data risk assessment, define the scope and scoring criteria, inventory and classify the data, map its flows and access, identify threats, evaluate controls, score each risk, and assign treatment.
The steps below follow the prepare, conduct, communicate, and maintain cycle described in NIST guidance.
Scope names the business units, systems, data types, environments, and regulations the assessment covers. Agree on the likelihood scale, impact scale, and risk appetite before discovery starts, so scores stay consistent across every team that contributes findings.
Discovery scans structured databases, file shares, SaaS applications, cloud object storage, backups, and logs for sensitive content. Record each asset with its location, business owner, technical custodian, approximate record count, and retention period.
Every asset without a named owner counts as a finding in its own right. Unowned data has no one accountable for its access reviews, retention, or deletion.
Classification labels each asset by the harm its exposure causes and by the rules that govern it. Many programs use tiers such as Public, Internal, Confidential, and Restricted.
Tag sensitive data types like payment card data, health records, credentials, and personally identifiable information (PII).
Flow mapping traces each Restricted or Confidential data set from collection through processing, storage, sharing, and deletion. Access mapping lists every human account, service account, API key, and vendor integration that reads or writes the data.
Excessive permissions surface during access mapping, such as a service account with write access to a production customer database that only one nightly job uses.
Threat identification lists the realistic ways each data set gets exposed: stolen credentials, cloud misconfiguration, insider misuse, vendor compromise, ransomware, and accidental sharing.
A structured threat analysis ties each threat to a specific weakness, such as a public storage bucket or a login without MFA.
Chained weaknesses matter as much as individual findings in a data risk assessment. Mapping the attack path from an exposed credential to a data store shows which low-severity findings combine into a high-severity risk.
Control evaluation checks whether encryption, access controls, logging, backups, data loss prevention, and retention policies work as documented. Test the controls instead of trusting the policy, because an enabled setting and an effective control are different things.
Scoring assigns each risk a likelihood value and an impact value on the agreed scales, then multiplies them into a severity score. The scoring model section below defines the scales, severity bands, and a worked example.
Treatment follows one of 4 options for each risk:
Rescore every treated risk to calculate residual risk, and record who approved any risk left above the appetite threshold.
Data, access, and cloud configurations change daily, so a point-in-time assessment starts aging on the day it closes. Track the treatment plan to completion, monitor for new data stores and permission changes, and rerun scoring when a trigger event occurs.
Data risk is calculated by multiplying the likelihood of an exposure by the impact of that exposure, with each factor rated on a defined scale. A 5-point scale for each factor produces scores from 1 to 25.
Risk score = Likelihood × Impact
A common banding treats scores of 1 to 4 as Low, 5 to 9 as Medium, 10 to 16 as High, and 20 to 25 as Critical. Each organization sets its own bands and response times as part of its risk appetite.
Inherent risk measures exposure before existing controls are considered, and residual risk measures it after them. Reporting both shows leadership how much risk the current controls remove and how much remains for treatment or acceptance.
The 5-by-5 model above counts as semi-quantitative, since it attaches numbers to descriptive ratings. NIST SP 800-30 Revision 1 describes qualitative, semi-quantitative, and quantitative assessment approaches, and it supplies example scales for likelihood and impact.
Quantitative methods such as FAIR (Factor Analysis of Information Risk) estimate risk in financial terms, using loss event frequency and loss magnitude. Boards and cyber insurers respond well to dollar figures, while a 5-by-5 matrix is faster to apply across hundreds of data assets.
A SaaS company finds a cloud database holding 500,000 customer records. Encryption at rest is disabled, 25 employees and one over-privileged service account have access, and an old backup of the database is stored in publicly readable storage.
The public backup makes exposure almost certain, so likelihood scores 5. Regulated personal data at that volume scores 5 for impact. The inherent risk score is 25, which lands in the Critical band.
Treatment deletes the public backup, enables encryption, cuts access to 6 named roles, and rotates the service account credentials. Likelihood drops to 2, but impact stays at 5 because the records still exist, leaving a residual score of 10 in the High band.
Tokenizing payment and national ID fields lowers impact to 3 and residual risk to 6, a Medium score the risk owner accepts within appetite.
A data risk assessment differs from a DPIA and other assessments in whose risk it measures and what triggers it. A DRA measures risk to the organization's data, while a DPIA measures risk to the people the data describes.
Most organizations run several of these on the same data. A DRA identifies the high-risk data sets, and a DPIA then examines the specific processing activities that use them.
Frameworks define how to run a data risk assessment, while regulations define when one is mandatory and what evidence it produces.
Data risk assessments miss the highest risks when they cover only sanctioned systems. Cloud storage, shadow data, forgotten copies, vendors, and AI tools hold sensitive data outside the systems most inventories list.

Cloud object storage, managed databases, and snapshots change configuration through infrastructure code, console edits, and automation. A bucket that was private during the assessment turns public after one policy change.
Include internet-facing checks in the assessment, because internal scans do not show what the internet sees. External attack surface management shows which data stores and services are reachable from the internet, the same view an attacker scans.
Shadow data refers to sensitive data copied into places no inventory tracks: SaaS exports, test databases, analytics sandboxes, and collaboration workspaces. Developer tools are a frequent source, since API collections and scripts hold real tokens and sample records.
A year-long CloudSEK investigation found more than 30,000 publicly accessible Postman workspaces leaking access tokens, refresh tokens, third-party API keys, business data, and customer PII. None of those workspaces belonged to a system that a typical data inventory lists.
Old backups, retired application databases, verbose logs, and former employees' file shares make up dark data, meaning information an organization stores but no longer uses or governs.
Assign an owner to each dark data store during the assessment, then delete, archive, or bring it under the same controls as live data.
Vendors that host, process, or support sensitive data extend the assessment boundary.
Score each vendor by the data it reaches, its access level, where it stores copies, and its demonstrated controls, using the same likelihood and impact scales as internal systems.
A breach at a shared cloud provider reaches every customer tenant at the same time.
In March 2025, CloudSEK reported a threat actor selling 6 million records allegedly exfiltrated from Oracle Cloud's SSO and LDAP systems, affecting more than 140,000 tenants, a claim Oracle publicly denied. Continuous vendor risk monitoring and planning for a third-party data breach keep provider risk visible between annual reviews.
AI adoption creates data stores that older assessments never listed: prompt histories, retrieval-augmented generation (RAG) indexes, vector databases, fine-tuning datasets, and AI agent permissions.
Employees using unsanctioned tools, known as shadow AI, send sensitive data to providers with unknown retention terms.
Score AI data stores on the same likelihood and impact scales as any other data asset. IBM's 2026 Cost of a Data Breach Report put the global average cost of an AI model inversion attack at $6 million, which reflects how hard training data is to protect once a model exposes it. Leaked AI API keys belong in the same review, because a stolen key can read uploaded files and stored prompts.
Organizations run a data risk assessment at least once a year, continuously in fast-changing cloud environments, and again whenever a trigger event changes the data or its exposure.
Regulated schedules set the minimum frequency, not the ideal one. PCI DSS reviews targeted risk analyses at least every 12 months, and DPDP Rule 13 sets a 12-month cycle for Significant Data Fiduciaries.
Data risk assessment tools automate discovery, classification, access analysis, and scoring, so the inventory stays current between formal reviews.
How long does a data risk assessment take?
It varies with scope. A single application takes days, while a first enterprise-wide assessment across cloud, SaaS, and on-premises systems takes weeks to months.
Who should perform a data risk assessment?
Security or data governance teams lead it, with data owners, privacy, legal, and IT supplying input. DPDP Rule 13 requires an independent person for Significant Data Fiduciary assessments.
Is a data risk assessment the same as a data audit?
No. A data audit checks compliance against defined requirements, while a data risk assessment estimates the likelihood and impact of future data exposure.
Can a small business run a data risk assessment without DSPM tools?
Yes. A spreadsheet inventory, a 5-by-5 scoring matrix, and access reviews in each SaaS admin console cover a small environment.
Does a data risk assessment cover paper records?
Yes. Paper records containing personal or health data fall under GDPR filing-system rules and HIPAA, so physical storage and disposal belong in scope.
What does a data risk assessment template include?
A data risk assessment template includes inventory fields, classification tiers, likelihood and impact scales, severity bands, risk register columns, and a treatment plan section.
