PHI and AI: Where Protected Health Information Actually Leaks
Everyone knows not to paste a chart into a chatbot. The leaks happen everywhere else.
What Counts as PHI in an AI Context
Nobody in your organization is copying a full medical record into a chatbot. That scenario, the one every warning email describes, is the least common way PHI reaches an AI tool. The common ways look like work: a scheduler summarizing a voicemail, a case manager drafting a prior-auth appeal, a meeting assistant sitting in on a care coordination huddle. Each one feels like a task, not a disclosure.
Protected health information is any individually identifiable health information held or transmitted by a covered entity or business associate, in any form. In an AI context, the definition does not shrink, but the surface expands. A prompt is a transmission, a generated summary is a new copy, and a meeting transcript is a record. The eighteen HIPAA identifiers still define "identifiable," and the practical test for a prompt is simple: if a name, date, location, or detail in it could point back to one patient, PHI just moved.
The part staff miss is that identifiability survives paraphrase. "78-year-old woman on the cardiac unit who fell twice this week" names no one and identifies someone.
The Six Places PHI Actually Reaches AI
Prompts, paraphrased
Not pasted charts, retyped fragments. The patient described from memory, the lab value quoted in a question. Small enough to feel harmless, identifiable enough to count.
Meeting assistants and transcripts
A notetaker bot in a huddle, an ambient recorder in a hallway conversation. The tool is the third party in the room, and several popular ones train on what they hear on lower tiers.
Documents sent for summarizing
Referral PDFs, discharge instructions, appeal letters run through a document AI. The whole file transmits, not just the part the user cared about.
Calendar and email metadata
Meeting titles like "Care conference: [patient name]" and inboxes an AI assistant can read. Assistants with inbox and calendar awareness now arrive by default in productivity suites.
Generated output
A summary containing PHI is new PHI, living wherever the output landed: a chat history, a personal note app, a downloads folder on a home device.
Uploads to build something
Spreadsheets fed to an analysis tool, datasets given to a copilot to "just clean up." The builder's version of the paraphrase problem, at larger volume.
The Control That Matches Each Leak
The pattern in that list: PHI leaks at surfaces policy cannot see. The matching controls, in the order they pay off: a sanctioned, governed AI path so the work has somewhere safe to happen, PHI screening between users and models so the paraphrase problem is caught mechanically rather than by memory, tool rules that name meeting assistants and document AI specifically, calendar and title hygiene in scheduling workflows, and retention rules for AI output, which is the copy everyone forgets exists.
Training still matters, but note what it is for, teaching staff to recognize the six surfaces, not to memorize the eighteen identifiers. Nobody fails HIPAA by forgetting identifier twelve. They fail it by not recognizing a transmission as one.
What De-Identification Does and Does Not Buy
Properly de-identified data is not PHI, and HIPAA has two recognized methods, expert determination and safe harbor removal of the eighteen identifiers. Two cautions apply in an AI context. First, casual redaction is not de-identification; removing the name while keeping the story usually fails the test. Second, several AI vendors reserve rights to use de-identified data, which is legal under HIPAA and still worth a contractual look if your data governance standards are stricter than the statute.
PHI and AI: Common Questions
Is it a HIPAA violation to put PHI into ChatGPT?
Into consumer ChatGPT, yes: no BAA exists, so the disclosure is improper regardless of intent. Into a BAA-covered enterprise deployment with proper controls, PHI can be permissible. The tier is the difference.
Does removing the patient's name make a prompt safe?
Not by itself. Identifiability is about the combination of details. Dates, locations, rare conditions, and unit references can identify a patient with no name attached.
Is a meeting transcript PHI?
If the meeting discussed identifiable patients, yes, and so is the AI-generated summary of it. Both need the same protections as any other record.
What should we do first if PHI already reached an unsanctioned AI tool?
Treat it as a potential breach. What was disclosed, to which vendor and tier, and whether the four-factor assessment requires notification. Then fix the surface it leaked through, because the same surface will leak again.
Does a BAA make PHI in AI safe?
It makes it lawful to transmit. Safe requires the rest, access controls, screening, audit trails, and staff who recognize the six surfaces.
Related Reading
Continue across the HIPAA and AI governance cluster
HIPAA & AI Compliance
How HIPAA applies to AI tools, what OCR expects, and how to achieve compliance without blocking innovation
Read article → Case StudyThe Samsung ChatGPT Incident
How Samsung's employees exposed sensitive data and what healthcare can learn from it
Read article → Vendor ComparisonHIPAA-Compliant AI Platforms: The Independent 2026 Comparison
Five product categories claim the same label. An independent comparison of what each one actually does.
Read article →The First Control Is a Policy Staff Can Follow
Generate a healthcare-ready draft in minutes, then decide which tools earn a place in it.