Back to Blog
|18 min readDeveloper Guide

Five Healthcare Compliance Violations We Found Everywhere

Real code patterns from 3,000+ healthcare repositories, what they violate, and how to fix them.

By Satish Singh, CEO, Scrutora

We scanned over 3,000 healthcare repositories for code-level compliance violations. The findings weren’t exotic. They were patterns every developer has seen, just never framed in the context of regulatory compliance.

This is a developer’s guide to five patterns we found repeatedly, across organizations including Google, CDC, NIH, the NHS, and the most widely deployed open-source EMR platforms in the world. Each pattern has appeared in healthcare code processing real patient data. Each one is fixable in minutes.

If you write code that touches health data, search your codebase for these before your next audit does.

1. PHI in Application Logs

What it looks like

Python
logger.info(f"Patient registered: {patient.name} {patient.aadhaar_number}")
logger.debug(f"Processing record: {patient}")
log.error(f"Failed to update patient: {patient_data}")

Why it’s a violation

Application logs are not protected storage. They flow to stdout, get captured by log aggregators (Splunk, ELK, CloudWatch, Datadog), get shipped to SaaS vendors, end up in backup archives, and get reviewed by operations teams who have no BAA.

The moment patient data is written to a log, it becomes PHI in an unprotected channel. HIPAA §164.530(j) requires documentation of all PHI disclosures. GDPR Article 5(1)(f) requires appropriate security. A plaintext log of patient names, IDs, or diagnoses violates both.

Found in: DIVOC (India’s vaccination certificate system), where Aadhaar numbers, full names, DOB, gender, phone, and addresses are written to stdout for every certificate generated. ABDM hip-service (India’s national health data exchange), where patient name, ID, gender, and address are logged to ELK stack. Multiple EMR systems in the US and Europe.

The fix

Never log raw patient data. If you need to log for debugging, log references, not content.

Log references, not contentPython
# Bad
logger.info(f"Processing patient: {patient.name} {patient.ssn}")

# Good
logger.info(f"Processing patient_id={patient.id} operation=update")
Field-level redaction middlewarePython
SENSITIVE_FIELDS = {'ssn', 'name', 'dob', 'aadhaar', 'phone', 'email', 'address'}

def redact(record):
    if isinstance(record, dict):
        return {k: '***' if k in SENSITIVE_FIELDS else v
                for k, v in record.items()}
    return record
Search your codebase nowBash
grep -rn "logger\.\(info\|debug\|error\).*patient" --include="*.py"
grep -rn "console\.log.*patient" --include="*.js" --include="*.ts"
grep -rn "log\.\(info\|debug\|error\).*patient" --include="*.java"

You’ll probably find something.

2. TLS Verification Disabled in Production

What it looks like

Python
# TODO: Fix before prod
response = requests.get(url, verify=False)
JavaScript
const agent = new https.Agent({ rejectUnauthorized: false });
Java
TrustStrategy trustAll = (chain, authType) -> true;  // trust everything

Why it’s a violation

Disabling TLS verification means your code accepts any certificate, including ones presented by man-in-the-middle attackers. HIPAA §164.312(e)(1) requires protection of ePHI during transmission. GDPR Article 32 requires appropriate technical measures during data transfer. Neither allows “we trust every certificate.”

The TODO comment doesn’t make it acceptable. The TODO comment proves someone knew it was wrong and shipped it anyway.

Found in: VA notification-api (Python), verify=False with a #nosec suppression comment in production code handling veteran medical notifications. NCI NBIA (Java), TLS verification disabled via a TrustStrategy that returns true for every certificate. ProjectEKA Jataayu (Java, Android), SSL verification completely disabled in India’s national health records mobile app. IBM LinuxForHealth (Python), multiple verify=False calls in bulk FHIR export handling.

The fix

Proper TLS handlingPython
# Bad
response = requests.get(url, verify=False)

# Good (for self-signed certs in development)
response = requests.get(url, verify='/path/to/ca-cert.pem')

# Good (for production: uses system CA bundle)
response = requests.get(url)
Search your codebaseBash
grep -rn "verify=False" --include="*.py"
grep -rn "rejectUnauthorized.*false" --include="*.js" --include="*.ts"
grep -rn "TrustStrategy\|ALLOW_ALL_HOSTNAME_VERIFIER" --include="*.java"

Look for TODOs near TLS code. Fix them today, not “before prod.”

3. Math.random() for Security Operations

What it looks like

JavaScript
const otp = Math.floor(1000 + Math.random() * 9000);
const sessionId = Math.random().toString(36).substring(2);
const token = Date.now() + '-' + Math.random();
Python
import random
salt_position = random.randint(0, len(value))
api_key = str(random.random())

Why it’s a violation

Math.random() and Python’s random module are not cryptographically secure. They use deterministic algorithms (Mersenne Twister, LCG) that can be predicted with a few observed outputs. When you use them for security-sensitive operations like OTPs, session IDs, password salts, or cryptographic tokens, you’ve built security theater.

This violates HIPAA §164.312(a)(2)(iv) (encryption and decryption). HIPAA requires cryptographic protection of ePHI. Predictable randomness is not protection.

Found in: CDAC MyHealthRecord (India’s first national Personal Health Record system), Math.random() for 4-digit OTP generation used to verify Aadhaar for account access. OHDSI Data2Evidence (clinical research platform), Math.random() for cryptographic salt positioning in credential processor. Multiple healthcare repos using Math.random() for session tokens and API keys.

The fix

Use crypto APIsJavaScript
// Bad
const otp = Math.floor(1000 + Math.random() * 9000);

// Good (Node.js)
const crypto = require('crypto');
const otp = crypto.randomInt(1000, 10000);

// Good (Browser)
const otp = crypto.getRandomValues(new Uint32Array(1))[0] % 9000 + 1000;
Use secrets modulePython
# Bad
import random
otp = random.randint(1000, 9999)

# Good
import secrets
otp = secrets.randbelow(9000) + 1000
Search your codebaseBash
grep -rn "Math\.random" --include="*.js" --include="*.ts"
grep -rn "random\.random\|random\.randint" --include="*.py"
grep -rn "new Random()" --include="*.java"

4. Patient Data Sent to External AI APIs

What it looks like

Python
response = openai.chat.completions.create(
    model="gpt-4",
    messages=[{"role": "user",
               "content": f"Analyze this patient record: {patient_data}"}]
)

Why it’s a violation

Every API call to a third-party LLM sends your data to that provider’s servers. The moment patient data leaves your environment, you need a Business Associate Agreement (BAA) with the receiving party. OpenAI’s standard API does not include a BAA by default. Neither does Anthropic’s, Google’s Gemini, or most other LLM providers.

Without a BAA, sending PHI to an LLM is a HIPAA §164.308(b)(1) violation. The data is now on a third-party server with no contractual protection, no breach notification obligation, and no access controls you can audit.

Found in: TeleICU middleware (rural ICU monitoring in India), patient monitor images sent to OpenAI’s chat completions API. Multiple funded healthcare startups sending patient notes, symptoms, and medical history to LLMs without enterprise agreements.

The fix

1. Sign a BAA with your LLM provider. OpenAI offers this for Enterprise tier. Anthropic offers this for Claude. Azure OpenAI offers BAAs.

2. De-identify before sending. Strip all 18 HIPAA identifiers before the API call, not after.

3. Use on-device or private deployments. Smaller models (Llama, Mistral, Phi) running on your own infrastructure don’t leave your environment.

Progressive improvementPython
# Bad
def analyze_patient(patient_record):
    return openai.chat.completions.create(
        messages=[{"role": "user", "content": str(patient_record)}]
    )

# Better
def analyze_patient(patient_record):
    deidentified = remove_phi(patient_record)  # strip names, dates, IDs
    return openai.chat.completions.create(
        messages=[{"role": "user", "content": str(deidentified)}]
    )

# Best (with BAA in place)
def analyze_patient(patient_record):
    return azure_openai.chat.completions.create(  # BAA-covered deployment
        messages=[{"role": "user", "content": str(patient_record)}]
    )

5. XXE Vulnerabilities in Medical Data Parsing

What it looks like

Java
DocumentBuilderFactory factory = DocumentBuilderFactory.newInstance();
DocumentBuilder builder = factory.newDocumentBuilder();
Document doc = builder.parse(inputStream);
Python
from lxml import etree
tree = etree.parse(xml_file)  # default resolver enabled

Why it’s a violation

XML External Entity (XXE) attacks allow an attacker to include arbitrary server-side files, make outbound network requests, or exfiltrate data by crafting malicious XML documents. Healthcare systems parse XML constantly: HL7 messages, DICOM metadata, FHIR XML, CCD/CCDA documents, lab results.

The default configuration of most XML parsers in Java and Python is vulnerable. You have to explicitly disable external entities. Most healthcare code we scanned didn’t.

Found in: i2b2 (Harvard clinical research platform, used at 200+ academic medical centers), 18 XXE vulnerabilities. NCI NBIA (NIH cancer imaging archive), 7 XXE vulnerabilities. OpenClinica (world’s most widely used clinical trials management system), 24 XXE vulnerabilities. CommCare (WHO/UNICEF platform, 500K+ health workers in 80+ countries), 19 XXE vulnerabilities.

The fix

Disable external entitiesJava
// Bad
DocumentBuilderFactory factory = DocumentBuilderFactory.newInstance();

// Good
DocumentBuilderFactory factory = DocumentBuilderFactory.newInstance();
factory.setFeature(
    "http://apache.org/xml/features/disallow-doctype-decl", true);
factory.setFeature(
    "http://xml.org/sax/features/external-general-entities", false);
factory.setFeature(
    "http://xml.org/sax/features/external-parameter-entities", false);
factory.setFeature(
    "http://apache.org/xml/features/nonvalidating/load-external-dtd", false);
factory.setXIncludeAware(false);
factory.setExpandEntityReferences(false);
Use defusedxmlPython
# Bad
from lxml import etree
tree = etree.parse(xml_file)

# Good
from lxml import etree
parser = etree.XMLParser(resolve_entities=False, no_network=True)
tree = etree.parse(xml_file, parser)

# Best
from defusedxml.lxml import parse
tree = parse(xml_file)
Search your codebaseBash
grep -rn "DocumentBuilderFactory\|SAXParserFactory\|XMLInputFactory" --include="*.java"
grep -rn "lxml\|xml\.etree\|xml\.sax" --include="*.py"

The Pattern Behind the Patterns

None of these five findings are clever attacks. They’re not zero-days. They’re not sophisticated. They’re default configurations that shipped to production because nobody checked.

That’s the real story. Compliance isn’t failing because teams don’t care. It’s failing because there’s no layer of review that catches code-level violations. Policy reviews catch policies. Infrastructure audits catch infrastructure. Security scanners catch CVEs. But nothing systematically checks whether your code logs a patient SSN, sends patient data to OpenAI, or disables TLS with a TODO comment.

That’s the gap.

Your certificate says you’re HIPAA-compliant. Your code says otherwise.

Try It Yourself

Run the grep commands above on your own codebase. You’ll find something. Then fix it.

If you want a systematic scan, you can run Scrutora against your repo for free at scrutora.com. No signup for public repositories. HIPAA, GDPR, SOC 2, DPDPA, and six other frameworks supported.

Published by Satish Singh, Founder and CEO of Scrutora. This post draws on findings from the Healthcare Code Compliance Security Index 2026, a scan of over 3,000 healthcare repositories globally.

Scan Your Codebase for These Patterns

Scrutora checks for all five violation patterns above , plus 700+ more, across HIPAA, GDPR, SOC 2, DPDPA, PCI DSS, and five other frameworks.