🚀 CloudSEK featured in the 2026 Frost Radar™!
Read more
Sensitive data is information that causes financial, legal, physical, or reputational harm to a person or organization when it is accessed, disclosed, altered, or destroyed without authorization.
Government ID numbers, health records, payment card details, passwords, API keys, source code, and trade secrets are all forms of sensitive data.
Personal information makes up only one part of the sensitive data an organization holds. A database password identifies no individual, yet an attacker holding that password reaches every customer record stored behind it. Legal definitions add a second layer, because GDPR, the CCPA, HIPAA, and India's DPDP Act each protect a different list of categories.
Attackers pursue sensitive data for the access and income it brings them. Stolen credentials open systems, stolen health and financial records feed fraud, and stolen corporate files become material for extortion.
Information qualifies as sensitive data when its unauthorized exposure produces measurable harm, grants access to protected systems, or places it in a legally protected category. Security and privacy teams apply the following tests before assigning a sensitivity label.
Data sensitivity changes across the life of a record instead of staying fixed at the point of collection. Merger plans lose their sensitivity on the day of public announcement. A former employee's medical leave file stays sensitive for as long as the organization keeps it.
Sensitive data divides into personal identifiers, special category personal data, financial data, authentication secrets, intellectual property, security data, and government information. Each type carries a distinct harm profile and attracts a distinct group of attackers.

Direct identifiers single out one person on their own, for example, a passport, government IDs such as Aadhaar, and Social Security numbers. Indirect identifiers, such as a home address, precise geolocation, and a device ID, become identifying once linked to a name or account.
Identity thieves combine both kinds to open accounts and pass verification checks in a victim's name, a pattern covered in CloudSEK's guide to personally identifiable information (PII).
Medical records, genetic test results, health history, treatment details, and biometric templates are classified as Protected Health Information (PHI). It can expose someone to discrimination, blackmail, or even physical harm. Similarly, religious beliefs, sexual orientation, political opinions, and trade union membership fall into this category due to their potential impact on privacy and personal safety.
Privacy laws in the EU, California, and other jurisdictions apply stricter processing rules to this type.
Payment card numbers, bank account numbers, and transaction histories give criminals a direct route to money. Card verification codes and PINs raise the risk further, because a fraudster combines them with a card number to complete purchases.
Credit reports, tax filings, and salary records complete this type.
Credentials and secrets describe no one, but they open systems. Passwords, OAuth refresh tokens, and cloud access keys belong here, alongside SSH private keys and MFA recovery codes.
A single exposed cloud access key reaches every storage bucket and database that its assigned role permits.
Source code, product designs, and proprietary algorithms represent years of research investment. Unreleased financial results, acquisition plans, and negotiated customer pricing fit this type. Exposure causes permanent damage, since a competitor cannot unlearn a stolen design.
Network diagrams, firewall rules, and vulnerability scan reports document how an organization defends itself. Incident response records and penetration test results list known weaknesses.
In an attacker's hands, this material works as a map for plotting an attack path toward more valuable data.
Governments classify information by the damage its disclosure causes to national security. In the United States, Executive Order 13526 sets 3 classification levels: Confidential, Secret, and Top Secret.
Citizen records, including tax filings, voter rolls, and benefits data, stay sensitive outside that classification system.
The difference between sensitive data, personal data, and PII lies in what each term measures. Personal data and PII measure whether information relates to an identifiable person, while sensitive data measures how much harm unauthorized exposure causes.
Personal data and sensitive data overlap only in part, so neither term replaces the other. A customer's shipping address is personal data with low sensitivity. Source code identifies no one, yet it counts as high-value sensitive data for any software company.
Sensitive personal data covers the overlap between the two ideas. It refers to personal data that a specific law singles out for stricter handling, such as GDPR special categories or California's sensitive personal information.
Classification assigns each sensitive data set a level that dictates how it is accessed, stored, shared, retained, and destroyed. Many corporate security policies use 4 levels: Public, Internal, Confidential, and Restricted.
US federal agencies rate sensitivity on a separate impact scale defined in FIPS 199. That standard rates the potential impact of a loss of confidentiality, integrity, or availability as low, moderate, or high. The rating determines the baseline security controls an agency selects from NIST SP 800-53.
No single global legal definition of sensitive data exists. Each privacy law protects a fixed list of categories, and the lists differ by jurisdiction and sector.
GDPR Article 9 bans processing of special category data unless an exemption applies, such as explicit consent, employment law obligations, or medical diagnosis by a health professional.
The special categories listed in GDPR Article 9 cover racial or ethnic origin, political opinions, religious or philosophical beliefs, and trade union membership. The list extends to genetic data, biometric data used to uniquely identify a person, health data, and data about a person's sex life or sexual orientation.
Breaches of the Article 9 conditions carry fines of up to 20 million euros or 4% of worldwide annual turnover, whichever is higher.
California treats sensitive personal information as a subset of personal information, and consumers hold a right to limit how businesses use and disclose it.
The definition covers government identifiers such as Social Security, driver's license, and passport numbers; account logins with passwords; financial account and card numbers with access codes; and precise geolocation.
Racial or ethnic origin, religious beliefs, union membership, genetic and biometric data, health data, sex life, sexual orientation, and the contents of private mail, email, and text messages fall inside the definition. Amendments added citizenship and immigration status from January 1, 2024, and neural data from January 1, 2025.
HIPAA protects health information held by covered entities and their business associates. Covered entities include health plans, healthcare clearinghouses, and providers that conduct standard electronic transactions.
Protected health information (PHI) exists in electronic, paper, and oral form. The Safe Harbor method under the HIPAA Privacy Rule lists 18 identifiers, including names, medical record numbers, and full-face photographs, that are removed to de-identify a data set.
Payment card rules under PCI DSS separate cardholder data from sensitive authentication data. Cardholder data includes the primary account number (PAN), cardholder name, expiration date, and service code.
PCI DSS requires the PAN to be unreadable wherever it is stored. Sensitive authentication data covers full track or chip data, card verification codes, and PINs, and PCI DSS bans retaining it after authorization, even in encrypted form.
India's Digital Personal Data Protection Act, 2023 applies one framework to all digital personal data and creates no separate sensitive category.
The DPDP Rules, notified in November 2025, phase in the core obligations over 18 months, with most duties taking effect in May 2027.
Until that date, the 2011 SPDI Rules under Section 43A of the IT Act remain in force. Those rules define sensitive personal data or information as passwords, financial information, physical and mental health conditions, sexual orientation, medical records, and biometric information. Sensitivity still shapes the DPDP Act, since the government designates Significant Data Fiduciaries partly on the volume and sensitivity of the personal data they process.
A US Department of Justice rule codified at 28 CFR Part 202 took effect on April 8, 2025. It prohibits or restricts transactions that give countries of concern or covered persons access to bulk US sensitive personal data.
China (including Hong Kong and Macau), Cuba, Iran, North Korea, Russia, and Venezuela are the listed countries of concern.
Covered personal identifiers, precise geolocation data, biometric identifiers, human 'omic data, personal health data, and personal financial data make up the rule's sensitive personal data categories. Bulk thresholds range from more than 100 US persons for human genomic data to more than 100,000 US persons for covered personal identifiers.
Sensitive data gets exposed through stolen credentials, secrets published in code and developer tools, cloud misconfiguration, third-party compromise, insider mistakes, and uploads to AI tools. Several of these paths begin outside the systems that store the data.
Infostealer malware copies saved browser passwords, session cookies, and VPN logins from infected devices and sends them to criminal marketplaces. The 2024 Snowflake campaign showed the impact.
Mandiant tied data theft from about 165 Snowflake customer organizations to credentials stolen by infostealers, some dating to 2020, on accounts without multi-factor authentication, as reported by SecurityWeek. The investigations found no flaw in the Snowflake platform itself.
Security teams track leaked credentials across breach dumps and stealer logs to catch this exposure before an attacker logs in.
Developers place API keys, database connection strings, and cloud tokens in code during testing. Those secrets then reach public GitHub repositories, compiled mobile apps, and paste sites such as Pastebin. Anyone who finds a live key inherits its permissions.
API development platforms expose the same secrets when workspaces are set to public. A year-long CloudSEK investigation found more than 30,000 publicly accessible Postman workspaces leaking access tokens, refresh tokens, third-party API keys, business data, and customer PII.
AI provider credentials add a newer risk to the same exposure pattern. Leaked AI API keys let an attacker read uploaded files, stored prompts, and fine-tuned models, not just consume paid compute quota.
Public storage buckets, databases without authentication, and overly broad identity and access management (IAM) roles expose sensitive data without any intrusion. Internet-wide scanning services index exposed assets continuously, so attackers locate a newly exposed database without targeting the organization first.
Vendors that run payroll, host customer support, or process analytics keep copies of their clients' sensitive data. One compromised vendor exposes records from every client it serves, and the client that owns the data carries the notification duty.
A third-party data breach and a broader supply chain attack follow the same pattern, with the attacker entering through the least protected partner.
Employees expose sensitive data through misdirected emails, public file-sharing links, and personal cloud accounts used for work files. Malicious insiders act deliberately, copying customer lists or source code before resigning.
Staff paste customer records, contracts, and source code into AI assistants to work faster. In April 2023, Samsung engineers uploaded internal source code to ChatGPT, and Samsung restricted generative AI tools the next month, according to Bloomberg.
Shadow AI, meaning AI tools used without security approval, places that data outside any retention or access control.
IBM's 2026 Cost of a Data Breach Report found that more than 20% of organizations reported a breach targeting AI models or applications. Compromised APIs, apps, or plug-ins and cloud misconfigurations affecting AI workloads each accounted for 27% of those breaches.
A sensitive data leak is accidental exposure with no confirmed attacker, such as a publicly readable storage bucket. A data breach is unauthorized access to or theft of data, such as ransomware operators exfiltrating patient records.
Leaks turn into breaches the moment an attacker finds and uses the exposed data, which makes data leak detection a time-critical control.
Organizations face response costs, regulatory penalties, extortion, and lost business when sensitive data is exposed, and affected individuals face fraud and identity theft.
IBM's 2026 study put the global average cost of a data breach at $4.99 million, a 12% increase over 2025 and a record high. Breaches in which attackers used AI averaged $6 million.
In February 2024, attackers used compromised credentials to log in to a Change Healthcare Citrix portal that lacked multi-factor authentication. They moved through the network, stole data, and deployed ransomware 9 days after the first login. Change Healthcare later reported to the HHS Office for Civil Rights that approximately 192.7 million individuals were affected, according to Infosecurity Magazine.
Extortion groups now treat stolen data as leverage in its own right. Among ransomware incidents in IBM's 2026 study, attackers exploited brand reputation in 41% of cases, employee data in 35%, and intellectual property in 31%. Ransomware threat intelligence tracks which groups run leak sites and which sectors they target.
To protect sensitive data, organizations locate and classify it, restrict who reaches it, encrypt it, and watch for copies that surface outside the network.
Is an email address sensitive data?
No. An email address is personal data under GDPR, not special category data. It becomes sensitive once paired with a password or linked to health or financial records.
Are photographs considered sensitive data?
No, not by default. GDPR Recital 51 treats photographs as biometric special category data only when specific technical processing, such as facial recognition, uniquely identifies a person.
Is an IP address considered sensitive data?
No. The EU Court of Justice ruled in Breyer (2016) that dynamic IP addresses can be personal data, but IP addresses fall outside GDPR special categories.
Is encrypted data still sensitive data?
Yes. Encrypted personal data remains personal data under GDPR while a decryption key exists. GDPR Article 34 waives notice to individuals when breached data was rendered unintelligible through encryption.
Does anonymized data count as sensitive data?
No. GDPR Recital 26 places truly anonymized data outside the regulation. Pseudonymized data still counts as personal data, because additional information re-identifies the people behind it.
Are criminal records sensitive data under GDPR?
No, not as special category data. GDPR Article 10 governs criminal conviction and offense data separately and permits processing only under official authority or specific EU or member state law.
Who is responsible for protecting sensitive data in an organization?
The data controller, or the data fiduciary under India's DPDP Act, holds legal accountability for protecting sensitive data. Data owners assign classifications, and security teams operate the controls.
CloudSEK does not discover, classify, or encrypt sensitive data inside an organization's systems. Data security posture management, DLP, and key management tools cover that internal work.
Coverage from CloudSEK starts once sensitive data leaves the network. XVigil, CloudSEK's digital risk protection platform, monitors the surface, deep, and dark web for organization-specific exposure, including leaked credentials, customer records, and source code on paste sites, leaked-data marketplaces, and public code repositories.
A CloudSEK case study shows the pattern in practice. XVigil flagged a public GitHub repository, belonging to a user associated with a major Indian fintech company, that contained usernames, passwords, and database credentials. Those credentials offered attackers an initial access vector into internal infrastructure. The company closed the exposure by rotating the credentials and moving secrets into a secrets management tool.
