Skip to content
~/avishake
cd ../projects

avishake/phi-deidentification-pipeline

Clinical De-identification Pipeline.

Hospital archives → MedGemma training corpora, on-prem

AI / MLLLMsComputer Vision● active
Clinical De-identification Pipeline cover

model card · description

An on-premise pipeline that turns unstructured hospital archives into MedGemma training data without a single identifier leaving the building. Provider-agnostic multimodal page reading (Anthropic API, any OpenAI-compatible endpoint, or self-hosted vLLM), clinical/non-clinical classification, PHI removal from transcribed text and scanned images, automated leak verification, and AES-256 per-location packaging with SHA-256 manifests.

  • Image PHI removed with fractional rotated zonal masking; zero surviving identifiers on a 48-page clinical validation sample.
  • Data-driven redaction entities generated from extractor-reported PHI, eliminating O(patients × pages) matching at multi-terabyte scale.
  • In-place ZIP reading with header-only content sniffing surveyed an 11,631-file archive ~20× faster than full decompression.
  • Batched, checkpointed stages that resume after interruption.
  • Hardware and cost analysis (memory bandwidth, dense vs. MoE, per-page token economics) behind the on-prem decision.

evaluation

0
Leaked identifiers
~20× faster
Archive survey
11,631
Files surveyed
●now: shipping healthcare AI at Minion Technologies●latest commit: Tone Sphere (2 Oct 2026)●183★ across 33 open-source repos●5 publications · 10 citations●based in Kolkata, working with the world●press ~ to open the terminal●auto-synced 4 Oct 2026