We scanned over 3,000 healthcare repositories for code-level compliance violations. The findings weren’t exotic. They were patterns every developer has seen, just never framed in the context of regulatory compliance.
This is a developer’s guide to five patterns we found repeatedly, across organizations including Google, CDC, NIH, the NHS, and the most widely deployed open-source EMR platforms in the world. Each pattern has appeared in healthcare code processing real patient data. Each one is fixable in minutes.
If you write code that touches health data, search your codebase for these before your next audit does.
1. PHI in Application Logs
What it looks like
logger.info(f"Patient registered: {patient.name} {patient.aadhaar_number}")
logger.debug(f"Processing record: {patient}")
log.error(f"Failed to update patient: {patient_data}")Why it’s a violation
Application logs are not protected storage. They flow to stdout, get captured by log aggregators (Splunk, ELK, CloudWatch, Datadog), get shipped to SaaS vendors, end up in backup archives, and get reviewed by operations teams who have no BAA.
The moment patient data is written to a log, it becomes PHI in an unprotected channel. HIPAA §164.530(j) requires documentation of all PHI disclosures. GDPR Article 5(1)(f) requires appropriate security. A plaintext log of patient names, IDs, or diagnoses violates both.
The fix
Never log raw patient data. If you need to log for debugging, log references, not content.
# Bad
logger.info(f"Processing patient: {patient.name} {patient.ssn}")
# Good
logger.info(f"Processing patient_id={patient.id} operation=update")SENSITIVE_FIELDS = {'ssn', 'name', 'dob', 'aadhaar', 'phone', 'email', 'address'}
def redact(record):
if isinstance(record, dict):
return {k: '***' if k in SENSITIVE_FIELDS else v
for k, v in record.items()}
return recordgrep -rn "logger\.\(info\|debug\|error\).*patient" --include="*.py"
grep -rn "console\.log.*patient" --include="*.js" --include="*.ts"
grep -rn "log\.\(info\|debug\|error\).*patient" --include="*.java"You’ll probably find something.
2. TLS Verification Disabled in Production
What it looks like
# TODO: Fix before prod
response = requests.get(url, verify=False)const agent = new https.Agent({ rejectUnauthorized: false });TrustStrategy trustAll = (chain, authType) -> true; // trust everythingWhy it’s a violation
Disabling TLS verification means your code accepts any certificate, including ones presented by man-in-the-middle attackers. HIPAA §164.312(e)(1) requires protection of ePHI during transmission. GDPR Article 32 requires appropriate technical measures during data transfer. Neither allows “we trust every certificate.”
The TODO comment doesn’t make it acceptable. The TODO comment proves someone knew it was wrong and shipped it anyway.
The fix
# Bad
response = requests.get(url, verify=False)
# Good (for self-signed certs in development)
response = requests.get(url, verify='/path/to/ca-cert.pem')
# Good (for production: uses system CA bundle)
response = requests.get(url)grep -rn "verify=False" --include="*.py"
grep -rn "rejectUnauthorized.*false" --include="*.js" --include="*.ts"
grep -rn "TrustStrategy\|ALLOW_ALL_HOSTNAME_VERIFIER" --include="*.java"Look for TODOs near TLS code. Fix them today, not “before prod.”
3. Math.random() for Security Operations
What it looks like
const otp = Math.floor(1000 + Math.random() * 9000);
const sessionId = Math.random().toString(36).substring(2);
const token = Date.now() + '-' + Math.random();import random
salt_position = random.randint(0, len(value))
api_key = str(random.random())Why it’s a violation
Math.random() and Python’s random module are not cryptographically secure. They use deterministic algorithms (Mersenne Twister, LCG) that can be predicted with a few observed outputs. When you use them for security-sensitive operations like OTPs, session IDs, password salts, or cryptographic tokens, you’ve built security theater.
This violates HIPAA §164.312(a)(2)(iv) (encryption and decryption). HIPAA requires cryptographic protection of ePHI. Predictable randomness is not protection.
The fix
// Bad
const otp = Math.floor(1000 + Math.random() * 9000);
// Good (Node.js)
const crypto = require('crypto');
const otp = crypto.randomInt(1000, 10000);
// Good (Browser)
const otp = crypto.getRandomValues(new Uint32Array(1))[0] % 9000 + 1000;# Bad
import random
otp = random.randint(1000, 9999)
# Good
import secrets
otp = secrets.randbelow(9000) + 1000grep -rn "Math\.random" --include="*.js" --include="*.ts"
grep -rn "random\.random\|random\.randint" --include="*.py"
grep -rn "new Random()" --include="*.java"4. Patient Data Sent to External AI APIs
What it looks like
response = openai.chat.completions.create(
model="gpt-4",
messages=[{"role": "user",
"content": f"Analyze this patient record: {patient_data}"}]
)Why it’s a violation
Every API call to a third-party LLM sends your data to that provider’s servers. The moment patient data leaves your environment, you need a Business Associate Agreement (BAA) with the receiving party. OpenAI’s standard API does not include a BAA by default. Neither does Anthropic’s, Google’s Gemini, or most other LLM providers.
Without a BAA, sending PHI to an LLM is a HIPAA §164.308(b)(1) violation. The data is now on a third-party server with no contractual protection, no breach notification obligation, and no access controls you can audit.
The fix
1. Sign a BAA with your LLM provider. OpenAI offers this for Enterprise tier. Anthropic offers this for Claude. Azure OpenAI offers BAAs.
2. De-identify before sending. Strip all 18 HIPAA identifiers before the API call, not after.
3. Use on-device or private deployments. Smaller models (Llama, Mistral, Phi) running on your own infrastructure don’t leave your environment.
# Bad
def analyze_patient(patient_record):
return openai.chat.completions.create(
messages=[{"role": "user", "content": str(patient_record)}]
)
# Better
def analyze_patient(patient_record):
deidentified = remove_phi(patient_record) # strip names, dates, IDs
return openai.chat.completions.create(
messages=[{"role": "user", "content": str(deidentified)}]
)
# Best (with BAA in place)
def analyze_patient(patient_record):
return azure_openai.chat.completions.create( # BAA-covered deployment
messages=[{"role": "user", "content": str(patient_record)}]
)5. XXE Vulnerabilities in Medical Data Parsing
What it looks like
DocumentBuilderFactory factory = DocumentBuilderFactory.newInstance();
DocumentBuilder builder = factory.newDocumentBuilder();
Document doc = builder.parse(inputStream);from lxml import etree
tree = etree.parse(xml_file) # default resolver enabledWhy it’s a violation
XML External Entity (XXE) attacks allow an attacker to include arbitrary server-side files, make outbound network requests, or exfiltrate data by crafting malicious XML documents. Healthcare systems parse XML constantly: HL7 messages, DICOM metadata, FHIR XML, CCD/CCDA documents, lab results.
The default configuration of most XML parsers in Java and Python is vulnerable. You have to explicitly disable external entities. Most healthcare code we scanned didn’t.
The fix
// Bad
DocumentBuilderFactory factory = DocumentBuilderFactory.newInstance();
// Good
DocumentBuilderFactory factory = DocumentBuilderFactory.newInstance();
factory.setFeature(
"http://apache.org/xml/features/disallow-doctype-decl", true);
factory.setFeature(
"http://xml.org/sax/features/external-general-entities", false);
factory.setFeature(
"http://xml.org/sax/features/external-parameter-entities", false);
factory.setFeature(
"http://apache.org/xml/features/nonvalidating/load-external-dtd", false);
factory.setXIncludeAware(false);
factory.setExpandEntityReferences(false);# Bad
from lxml import etree
tree = etree.parse(xml_file)
# Good
from lxml import etree
parser = etree.XMLParser(resolve_entities=False, no_network=True)
tree = etree.parse(xml_file, parser)
# Best
from defusedxml.lxml import parse
tree = parse(xml_file)grep -rn "DocumentBuilderFactory\|SAXParserFactory\|XMLInputFactory" --include="*.java"
grep -rn "lxml\|xml\.etree\|xml\.sax" --include="*.py"The Pattern Behind the Patterns
None of these five findings are clever attacks. They’re not zero-days. They’re not sophisticated. They’re default configurations that shipped to production because nobody checked.
That’s the real story. Compliance isn’t failing because teams don’t care. It’s failing because there’s no layer of review that catches code-level violations. Policy reviews catch policies. Infrastructure audits catch infrastructure. Security scanners catch CVEs. But nothing systematically checks whether your code logs a patient SSN, sends patient data to OpenAI, or disables TLS with a TODO comment.
That’s the gap.
Your certificate says you’re HIPAA-compliant. Your code says otherwise.
Try It Yourself
Run the grep commands above on your own codebase. You’ll find something. Then fix it.
If you want a systematic scan, you can run Scrutora against your repo for free at scrutora.com. No signup for public repositories. HIPAA, GDPR, SOC 2, DPDPA, and six other frameworks supported.
Published by Satish Singh, Founder and CEO of Scrutora. This post draws on findings from the Healthcare Code Compliance Security Index 2026, a scan of over 3,000 healthcare repositories globally.
Scan Your Codebase for These Patterns
Scrutora checks for all five violation patterns above , plus 700+ more, across HIPAA, GDPR, SOC 2, DPDPA, PCI DSS, and five other frameworks.