Your Arch-Nemesis is a Data Scientist: What's the Difference Between Security and Privacy Work?
Aleatha Parker-Wood (Field CISO · Clearly AI)
BSidesSF 2026 · Day 1 · AMC Theatre 03
Overview
In a landscape increasingly dominated by data-driven decision-making and machine learning, the distinctions and overlaps between traditional security and privacy work have become critically important, yet often misunderstood. Aleatha Parker-Wood, a seasoned expert with extensive experience spanning both security R&D and privacy engineering, delivered a compelling talk at BSides SF, dissecting these differences from a practitioner's perspective. Her presentation, "Your Arch-Nemesis is a Data Scientist: What's the Difference Between Security and Privacy Work?", illuminated why security-centric approaches frequently fall short in addressing modern privacy challenges and how organizations can bridge this gap.
Key moments
- 0:40 Speaker's unique background in security and privacy
- 2:00 Fundamental legal rights granted by privacy laws
- 2:30 Why "personal data" is not just PII
- 4:30 Understanding purpose of processing and ongoing consent
- 5:40 Contrasting security and privacy failure scenarios
- 6:10 Your arch-nemesis: the well-intentioned data scientist
- 7:00 Security's confidentiality vs. privacy's appropriate data use
- 8:00 Granular data view and privacy's love for losing data
Your Arch-Nemesis is a Data Scientist: What's the Difference Between Security and Privacy Work?
Speakers: Aleatha Parker-Wood
Conference: BSides SF
YouTube: https://www.youtube.com/watch?v=tpQ5vWwZ8kU
Overview
In a landscape increasingly dominated by data-driven decision-making and machine learning, the distinctions and overlaps between traditional security and privacy work have become critically important, yet often misunderstood. Aleatha Parker-Wood, a seasoned expert with extensive experience spanning both security R&D and privacy engineering, delivered a compelling talk at BSides SF, dissecting these differences from a practitioner's perspective. Her presentation, "Your Arch-Nemesis is a Data Scientist: What's the Difference Between Security and Privacy Work?", illuminated why security-centric approaches frequently fall short in addressing modern privacy challenges and how organizations can bridge this gap.
Parker-Wood, currently a Field CISO at Clearly AI and formerly a Principal Privacy Engineer at Amazon, brings a unique blend of expertise, including a mid-career PhD in data governance, to this complex topic. She argues that security professionals, while adept at protecting data from external threats, often misinterpret the fundamental goals and adversaries of privacy. This talk is crucial for anyone involved in data handling, from engineers and data scientists to legal and executive teams, offering a pragmatic roadmap for building robust privacy programs that complement, rather than conflict with, existing security postures.
The core message is that privacy is not merely an extension of confidentiality but a distinct discipline focused on the appropriate use, retention, and deletion of personal data, driven by complex legal mandates. Understanding this shift in mindset – from battling external "bad guys" to managing the risks posed by well-intentioned internal data scientists – is paramount for effective data governance in today's regulated environment.
Background
▶ Watch: Speaker's unique background in security and privacy (0:40)
The prevailing understanding within many security circles is that privacy is primarily a facet of confidentiality, a component of the classic CIA triad (Confidentiality, Integrity, Availability). When asked about privacy, security professionals often default to discussions about cryptography, homomorphic encryption, or access control. While these are crucial for data protection, Parker-Wood emphatically states that this narrow view is a "horrible trap to fall into." Modern privacy laws, such as the General Data Protection Regulation (GDPR), California Consumer Privacy Act (CCPA), and California Privacy Rights Act (CPRA), do not even use the term Personally Identifiable Information (PII). Instead, they refer to "identifiable natural people" and "personal data," which encompasses anything that can identify an individual, either directly or indirectly. This includes not just obvious identifiers like social security numbers or driver's licenses, but also user IDs, session IDs, and any data linked to them.
The speaker highlights that privacy work is fundamentally different because it is "very much focused around what you do with personal data." It extends beyond merely who knows the data to encompass how it's used, where it's stored, and who it's shared with. This broader scope introduces a range of legal rights for individuals, including the right to know what data an organization holds about them (a Data Subject Access Request, or DAR), the right to correction, the right to deletion, and crucially, the right to revoke consent at any time. Furthermore, purpose limitation dictates that data can only be used for the specific purposes for which consent was originally obtained. This means organizations must have systems in place to check consent status "everywhere in your data handling infrastructure, every place, every time the data gets used."
The problem stems from a fundamental difference in mindset and the perceived "bad." For security, "bad" looks like a data breach, ransomware, or IP theft, often orchestrated by a malicious external actor. For privacy, "bad" looks like a failed deletion, a purpose limitation violation, or, "God forbid, got child data in my machine learning models." The primary adversary in privacy is not a nation-state hacker but "a very well-intentioned data scientist" who, in their eagerness to help the business, might inadvertently reidentify anonymous data, retain data beyond its legal purpose, or use data in ways not consented to. This necessitates a "kindler, gentler, more patient practice" of explaining risks rather than fighting off threats.
Key Findings
▶ Watch: Why "personal data" is not just PII (2:30)
Aleatha Parker-Wood's talk unveiled several critical distinctions and findings that challenge conventional security wisdom when applied to privacy:
- Personal Data vs. PII: The talk strongly emphasizes that most personal data does not contain PII as commonly understood by security professionals. Laws like GDPR and CCPA define personal data as anything that can identify a natural person, directly or indirectly. This includes not only direct identifiers but also user IDs, session IDs, or any data linked to them through foreign key relationships. This broader definition means security scanning tools designed to find specific PII patterns (e.g., social security numbers) are largely ineffective for identifying personal data in a privacy context.
- The Adversary is Internal and Well-Intentioned: A core finding is the shift in adversary. While security battles external "bad guys" and exploits, privacy's greatest challenge comes from internal, "very well-intentioned data scientists." These individuals, driven by business goals, are "most likely to reidentify data that was supposed to be anonymous," "keep things they shouldn't," or "not tell you what they have," all without malicious intent. This necessitates a focus on education, policy enforcement, and technical controls that prevent accidental misuse, rather than just malicious exploitation.
- Mindset Shift: From Data Retention to Data Deletion: Security is deeply concerned with data integrity, backups, and preventing data loss. Privacy, conversely, "loves losing data." The goal is to "delete all the data everywhere" and preferably "didn't collect it in the first place." This fundamental difference in philosophy — preserving data versus purging it — requires a complete re-evaluation of data retention policies and backup strategies. Absolute, irrevocable deletion is the privacy standard, with limited exceptions for tax, accounting, or fraud prevention data.
- Granular Data Understanding is Paramount: Security often categorizes data in "big buckets" (e.g., "critically sensitive data zone"). Privacy demands a "much more granular notion" of data handling. Organizations must understand the state of individual rows (e.g., minor, non-consenting user) and columns (e.g., biometrics, health data) within a dataset. This level of detail is necessary to apply purpose limitations and consent-based filtering, ensuring data is only used for permitted applications (e.g., preventing child data or biometrics from entering an ad targeting system).
- Legal Frameworks are Not Control Frameworks: Parker-Wood describes privacy laws as "more like a compiler spec" with "a huge amount of undefined behavior" that is only clarified through case law. Unlike mature security control frameworks (e.g., NIST 800-53), privacy control frameworks are "at least 10 years behind security," often developed with government use cases in mind, and don't map well to enterprise needs. This means compliance is not a "checkbox exercise" or achievable with an "out-of-the-box tool." Organizations must understand their unique legal obligations, including specific consent decrees from regulators, which are "laws that are just for you."
- Child Data as a "Glitterbomb": The handling of child data is highlighted as a critical area. "Once it has touched something, the thing that it has touched is forever tainted." Regulators impose severe penalties for mishandling child data. A compelling example cited is the FTC forcing Weight Watchers to not only pay a fine and delete data but also "delete their entire trained machine learning model and start over from scratch" because they couldn't prove child data hadn't been used in its training.
- Dynamic Consent and Its Operational Challenges: Unlike static data classifications, user consent can change "three times in a week." This means organizations cannot rely on cached consent statuses but "need to be able to check consent status every single time" data is used. This requirement introduces significant complexity for data pipelines and access control systems.
Technical Deep Dive
▶ Watch: Contrasting security and privacy failure scenarios (5:40)
Implementing effective privacy controls demands a significant technical overhaul, moving beyond traditional security paradigms. Parker-Wood outlined several key areas for a technical deep dive:
1. Upgrading the Data Catalog and Metadata:
The foundation of privacy work is an enhanced data catalog. Traditional security inventories typically identify systems and broad data types. For privacy, this needs to evolve to a much more granular level. Organizations must collect "a whole bunch of additional metadata" about their data assets. This involves sitting down with lawyers to understand specific legal obligations and translating those into technical metadata terms. The catalog must capture detailed semantics, labeling "table by table what is personal data" and, crucially, maintaining this information up to date. This often necessitates integration with code stacks or understanding data collected by vendors.
2. Granular Data Semantics and Purpose Modeling:
Privacy requires understanding not just that data is personal, but what kind of personal data it is and how it can be used. This means identifying:
- Row-level attributes: Is a user a minor? Have they consented? Have they revoked consent?
- Column-level attributes: Is this biometric data? Health data? Financial data?
- Purpose of Processing: What is this system used for? Is it for ads? Machine learning model training? Tax and accounting? Security logging?
The talk provides a concrete example: a dataset with two minors, one non-consenting user, and biometric data (e.g., face ID) for ads targeting. In some jurisdictions, biometrics cannot be used for ads. This means that out of an entire dataset, only a "little tiny bit of information" might be usable for a specific purpose. This necessitates "very fine-grained notions of access control and filtering."
3. Challenges with Automation (Scanners) and AI (LLMs):
- Scanners: While "everybody loves scanners" for automation, they are largely ineffective for privacy. Unlike PII (e.g., Social Security Numbers), personal data often involves generic identifiers like UUIDs (Universally Unique Identifiers), which are ubiquitous in systems. This makes it "very, very difficult to go and just scan and find personal data." Manual cataloging and "data hand annotated" by experts or lawyers are frequently required for critical data, following an 80/20 rule (automate what you can, manually verify the rest).
- LLMs: Despite the hype, Large Language Models (LLMs) "are not lawyers" and "should not be used to make decisions about the risk envelope of your business." They cannot recreate context if it wasn't captured during data ingestion, nor can they interpret opaque developer naming conventions. "Getting semantics wrong is expensive; it's usually worth the extra time and the extra money to go get a human judgment" from a lawyer for legal interpretations of personal data.
4. Access Control and Filtering Controls:
To enforce purpose limitations and consent, organizations need to "model purpose within your access control systems." This could involve:
- Role-Based Access Control (RBAC): Ensuring both human users and automated systems consume data using the correct role for its intended purpose.
- Attribute-Based Access Control (ABAC): Passing purpose as an attribute alongside other data characteristics to dynamically determine access.
Crucially, filtering controls must be integrated into the data pathway. This means actively scrubbing out "children," "non-consenting users," and "pesky biometric data" before it reaches systems where its use is prohibited. This ensures data scientists work with a "clean data set" that won't cause compliance issues.
5. Robust Data Lineage: User IDs and Timestamps:
A critical technical requirement, often overlooked by security, is the persistent association of user IDs and timestamps (when data was collected) with "everywhere the data goes." Parker-Wood warns, "You're really going to hate yourself later if you don't." Without this lineage, it becomes impossible to:
- Know which personal data belongs to which user.
- Track data retention periods.
- Fulfill Data Subject Access Requests (DARs) completely.
6. Absolute and Irrevocable Deletion:
Privacy demands "absolute 100% irrevocable deletion." This is far more stringent than typical security deletion practices:
- Hashing user IDs or delinking sessions is insufficient.
- Backups are problematic; they can only be kept for a limited "compliance SLA" (e.g., 15-60 days) unless the data falls under specific exceptions (tax, accounting, fraud, litigation holds). Even then, data retained for one purpose (e.g., tax) cannot be used for another (e.g., ads targeting) if the latter's retention period has expired.
- Deletion must be purpose-driven: "When a purpose expires, you can't use data for it." When all valid purposes expire, the data must be completely deleted.
7. Comprehensive Testing:
While privacy testing shares similarities with security testing (e.g., Canary users), it requires expansion. New types of Canary users are needed: "children," "consent canaries" (users who change consent status), and "deleted users." The most challenging aspect is "testing the completeness of your data subject access request," ensuring that all data belonging to a user is returned, which directly ties back to the necessity of robust user ID and timestamp lineage.
Demo / Proof of Concept
▶ Watch: Your arch-nemesis: the well-intentioned data scientist (6:10)
The talk by Aleatha Parker-Wood focused on a conceptual and architectural discussion of the differences between security and privacy work, providing practical tips and examples. It did not include a live demonstration or a specific proof of concept of a tool or system. Instead, the speaker used illustrative scenarios, such as the filtering of data for an ads targeting system based on user consent, age, and biometric data, to highlight the complexities and granular controls required for effective privacy management. These examples served to concretely explain the technical challenges rather than demonstrating a working prototype.
Defensive Implications
▶ Watch: Granular data view and privacy's love for losing data (8:00)
For organizations aiming to build robust defenses against privacy risks, the insights from this talk offer a critical roadmap for evolving beyond a purely security-centric approach:
- Re-evaluate and Upgrade Data Inventory: Defenders must move beyond simple PII identification. The data catalog needs a significant upgrade to capture granular metadata about all identifiable personal data, including its specific type (e.g., biometric, health), its associated purpose of processing, consent status, and data subject age. This necessitates close collaboration with legal teams to define what metadata is required to meet legal obligations and specific consent decrees.
- Shift Adversary Mindset and Educate Internally: Security teams must recognize that the primary adversary in privacy is often the well-intentioned internal data scientist. Instead of solely focusing on external threats, efforts should be directed towards educating internal teams about privacy risks, the legal implications of data misuse, and the profound business dangers (e.g., regulatory fines, reputational damage, forced deletion of ML models) associated with mishandling personal data. This requires a "kindler, gentler, more patient practice" of risk communication.
- Implement Granular Access and Filtering Controls: Traditional access controls are insufficient. Organizations need to develop and implement fine-grained access control systems that can model and enforce purpose limitations. This involves integrating filtering controls directly into data pipelines to automatically scrub out data (e.g., child data, non-consenting user data, certain biometric data) that is not permitted for a specific downstream use before it reaches processing systems. This ensures that data scientists and other consumers only receive data they are legally and ethically allowed to use.
- Prioritize Data Lineage and Identifier Propagation: A foundational defensive measure is to ensure that user IDs and timestamps of data collection are consistently propagated and linked to all personal data throughout its entire lifecycle, across all systems and transformations. This robust data lineage is absolutely critical for fulfilling Data Subject Access Requests (DARs), demonstrating compliance with data retention policies, and enabling complete and verifiable data deletions. Neglecting this will lead to significant compliance headaches and potential legal liabilities.
- Develop Absolute Data Deletion Capabilities: Defenders must engineer systems capable of absolute, 100% irrevocable deletion of personal data. This goes beyond logical deletion or hashing identifiers. It requires mechanisms to purge data from all active systems and, critically, from backups within defined compliance SLAs (e.g., 15-60 days), while respecting specific legal retention requirements for certain data types (e.g., tax, fraud logs). This is a complex engineering challenge that requires careful planning and implementation.
- Expand Privacy-Specific Testing: Security testing methodologies should be expanded to include privacy-specific scenarios. This means introducing new categories of Canary users, such as minors, users who have revoked consent, and users whose data has been requested for deletion. Rigorous testing for the completeness of Data Subject Access Requests is also paramount, verifying that all data linked to a user can be accurately identified, retrieved, and presented.
- Deepen Collaboration with Legal Counsel: Given the nascent and evolving nature of privacy laws and control frameworks, security and engineering teams must engage in much deeper and more frequent collaboration with legal counsel. Legal experts should guide the interpretation of regulations, the definition of metadata requirements, and the assessment of risk. Relying on automated tools or LLMs for legal judgment is explicitly cautioned against due to the high cost of semantic errors.
- Engage with the Privacy Practitioner Community: Recognizing that privacy control frameworks are still maturing, defenders should actively participate in and leverage the broader privacy practitioner community. This community is a valuable resource for sharing best practices, understanding emerging challenges, and collaboratively developing effective controls in an evolving regulatory landscape.
Key Takeaways
- Privacy is Distinct from Security: It's not just about confidentiality; privacy encompasses appropriate use, retention, and absolute deletion of all identifiable personal data, driven by legal rights and purpose limitations.
- The Adversary Shifted: In privacy, the primary risk often comes from well-intentioned internal data scientists who might inadvertently misuse data, rather than malicious external attackers.
- Granular Data Understanding is Critical: Effective privacy requires an extremely detailed data catalog that tracks personal data at the row and column level, including its purpose, consent status, and specific characteristics (e.g., child data, biometrics).
- Robust Data Lineage is Non-Negotiable: Propagating user IDs and timestamps with all personal data is essential for managing Data Subject Access Requests (DARs), enforcing retention policies, and enabling verifiable deletions.
- Laws are Complex, Not Checklists: Privacy laws are dynamic "compiler specs" with undefined behaviors, not simple control frameworks. Compliance requires deep understanding of unique legal obligations (including consent decrees) and cannot be achieved with off-the-shelf solutions.
- Deletion is Absolute and Purpose-Driven: Privacy demands 100% irrevocable data deletion, including purging from backups within short SLAs, and data can only be retained for specific, valid purposes; once a purpose expires, the data must be deleted for that purpose.
About the Speaker(s)
Aleatha Parker-Wood is a distinguished expert at the intersection of security, privacy, and machine learning. She is currently a Field CISO at Clearly AI, a company she describes as "amazing." Her extensive background includes five years as a Principal Privacy Engineer at Amazon, where she gained invaluable experience handling vast amounts of data. Prior to that, she spent five years running a security R&D lab for Symantec, showcasing her deep roots in cybersecurity. Parker-Wood also holds a mid-career PhD in data governance, further solidifying her expertise in the intricacies of data management and legal compliance. Her achievements are widely recognized, having been named one of the top women in machine learning for three consecutive years. She is also known for her ability to build high-powered technical teams and for her role as a "CEO whisperer about security," indicating her skill in translating complex technical concepts into strategic business implications.
Reviews
Dr. Zero (Offensive Security Researcher) — SOLID
Parker-Wood knows her material cold — the security-vs-privacy framing is genuinely useful, and the operational details (consent propagation, child-data contamination, purpose-limited deletion) are more concrete than most talks in this space. But this is a BSides explainer, not a research drop: it's well-executed education for practitioners who haven't thought hard about privacy engineering, not a novel contribution to anyone who has.
Heather Calloway (CISO) — SOLID
Parker-Wood knows this material cold and the core reframe — privacy's adversary is internal and well-intentioned, not external and malicious — is genuinely useful for security teams that have never built a privacy program. But this talk is aimed at practitioners building the pipes, not the executives and governance structures that decide whether those pipes get funded, prioritized, or enforced.