Home/Work/Ontora

Sorsix · internship capstone2026

Ontora

FHIR-native terminology mapping for laboratory data, with a confidence score that is measured, not assumed.

GLUClab A · abbreviation
FASTING BLOOD SUGAR, SER/PLASlab B · free text
L-002lab C · catalogue no.
LOINC 1558-6 fasting glucose, in a form any FHIR system can read
HIGH band
96.5%
top-1 precision
LOW band
8.6%
it knows when it is guessing
Calibration error
17.6%
published rather than hidden
Through Kafka
68rec/s
up from 42
Tests
250
in three languages

01Why it exists

A fasting glucose result is GLUC at one lab, FASTING BLOOD SUGAR, SER/PLAS at the next and L-002 at a third. Each is correct inside its own building and unreadable outside it. The result cannot follow the patient to another institution, national reporting counts one measurement three different ways, and a test that already happened gets ordered again because nobody downstream can tell. FHIR can carry that result anywhere, but only once it is coded as LOINC.

Ontora does that coding at scale. Records arrive over REST or Kafka. An embedding model proposes LOINC candidates, and every candidate is checked against a real FHIR terminology server. A reviewer decides, Ontora remembers the decision for that clinic, and publishes it as a standard FHIR ConceptMap.

The mapping is not the valuable part; a lookup table can map. What matters is that the system has been measured against hidden ground truth, so it knows which of its own answers to trust.
The review queue: records from three clinics as they arrived, their review status, and a live count of how many are eligible for one-click approval.
Queue Records from three clinics as they arrive, updated live. The banner counts the records whose top candidate is HIGH-confidence, verified by the terminology server and unflagged: the ones a reviewer can approve in one click.

02A confidence score you can check

The system was run over 660 evaluation records: 600 from three simulated clinics and 60 genuine de-identified hospital labs from MIMIC-IV. Each carries a ground-truth LOINC code that is never exposed through the API and never shown to the model. The 655 that have one are scored; the other five belong to lab items the 91-concept subset cannot honestly code.

Top-1 precision per confidence band, scored against hidden ground truth.

A cosine similarity is not a probability

The calibration chart shows it plainly. Expected calibration error is 17.6% over equal-count deciles, or 18.2% with equal-width bins. The shipped HIGH cut-off of 0.85 is deliberately conservative: the lowest cut-off that still holds 95% precision is 0.772, and the cost of every setting is tabulated by a script rather than typed by hand.

Observed top-1 accuracy against the cosine similarity that produced it, by equal-count decile, against a perfectly calibrated diagonal. Expected calibration error 17.6%.
Calibration Generated from the evaluation run by the project’s own script. The gap from the diagonal is the measurement, not a defect.

03Where it fails, on purpose

The same numbers, split by what each source sends:

That last row is the argument for the confidence spread. L-002 carries no meaning, so no model can recover it, and Ontora does not pretend otherwise: it scores those records LOW and sends them to a person. Once a person decides what L-002 means at St. Mary’s, the mapping memory applies that decision to every later L-002 from St. Mary’s without asking the model again, and the result still names the reviewer who made the original call as its author.

04Architecture

RESTPOST /api/records · X-Clinic-Key
Kafkaontora.records.v1 · 3 partitions, keyed by clinic
ontora-uiAngular 22 · REST + server-sent events
ontora-coreKotlin · Spring Boot 4
ingest · review · audit · FHIR
ontora-matcherPython · FastAPI · MiniLM · 91 LOINC concepts
HAPI FHIR JPAR4 terminology server · $validate-code
PostgreSQL 16records · suggestions · FHIR resources · audit
Why each service exists as its own thing
ServiceWhy it is its own thing
ontora-coreThe only component that writes. REST and Kafka are two adapters over one service, so both paths are audited identically.
ontora-matcherThe only part that needs PyTorch. Stateless, with 468 name vectors for 91 concepts in memory, so it scales and redeploys on its own schedule.
HAPI FHIR JPAA real terminology server, so validation is an external check rather than a list the system grades itself against.
PostgreSQL 16Five tables, FHIR resources as jsonb, eleven Flyway migrations. Hibernate only validates the schema; it never changes it.
KafkaA lab feed is a stream, not a request. Partitioning by clinic keeps each clinic’s results in order, and matching consumer concurrency to partitions is where 42 → 68 rec/s came from.

One record’s journey

  1. A clinic sends GLUC · 5.4 mmol/L · Patient/0042, over REST or Kafka.
  2. The core asks whether this clinic’s GLUC has been decided before.
  3. If it has, the mapping memory applies 1558-6 straight away and skips to the last step.
  4. Otherwise the matcher returns three candidates with their similarity.
  5. Every candidate is checked with $validate-code against HAPI FHIR, then gets a confidence band, a specimen check and a UCUM unit check.
  6. The reviewer’s queue updates live over server-sent events.
  7. The reviewer approves 1558-6; it is validated again at approval and remembered for this clinic.
  8. A FHIR Observation, Provenance and AuditEvent are stored. The reviewer’s identity comes from the JWT, never from the request body.

05The product

Amber means not yet portable, green means portable. Nothing on these screens describes a patient; they describe whether a result can be read somewhere else.

06Engineering decisions

Each of these had a real alternative.

  • 01 · MODEL CHOICE

    MiniLM over SapBERT

    SapBERT ranks 1.7 points better on top-1, but it is about five times slower, adds 420 MB, and makes the HIGH band less precise. That band carries the safety argument, so MiniLM stays the default.

  • 02 · TRANSACTIONS

    Publishing is its own action

    Publishing the ConceptMap is not part of approval. Otherwise a network call to HAPI sits inside the approval transaction, and approvals fail whenever the terminology server is down. The cost is one manual step.

  • 03 · CHECKS

    Warn, never block

    A candidate with a unit warning is right 2.5% of the time, but the check is a heuristic over free text, and a heuristic that blocks loses records silently. Flagged candidates are kept out of bulk approval, so a person always sees them.

  • 04 · SIGNALS

    The matcher never sees the axes

    It embeds names only; specimen and UCUM checks run afterwards in Kotlin. Two independent signals are useful precisely when they disagree, and folding them into one score would hide that.

  • 05 · TENANCY

    Memory is per clinic

    GLUC at one lab is a statement about that lab. Mappings are unique per clinic and source code, published under per-clinic code systems, and never shared across clinics.

  • 06 · SECURITY

    Duties split by the server

    Admin and reviewer roles never overlap: an admin gets a 403 on approve, a reviewer a 403 on the audit trail. Both are gated in the security filter chain, because method security throws where the catch-all handler would turn a 403 into a 500.

07Tests

JUnit 5 with Testcontainers and REST Assured for the core, against real PostgreSQL and Kafka containers; Vitest for the Angular UI; pytest with the real model for the matcher. Every FHIR resource Ontora produces (Observation, Provenance, AuditEvent, ConceptMap and the rest) is validated offline against the official R4 profiles, so a malformed resource fails the build instead of reaching an exchange partner. GitHub Actions runs the three suites as independent jobs.

08Stated limits, and what’s next

Where a measurement is weaker than it looks
  • The throughput figures were taken on a single laptop with every service running on it.
  • Accounts are held in memory, where a deployment would use an identity provider.
  • The live-update stream takes its token in the query string, because EventSource cannot set headers.
  • One clinic key is shared by every feed, where a deployment would issue one per feed.

What separates this from something that ships, in the order they unblock the most:

  1. Per-tenant terminology packs instead of one global LOINC subset.
  2. OIDC identity in place of in-memory accounts.
  3. A dead-letter topic and per-feed credentials. Today a malformed message is logged and dropped.
  4. Versioned mappings with effective dates, instead of overwriting in place.
  5. SNOMED CT as a second vocabulary: real work, since the specimen and unit checks do not carry over.
Kotlin 2.3Spring Boot 4.1Java 21 Angular 22Python 3.12FastAPI HAPI FHIR R4PostgreSQL 16Apache Kafka 4.2 Docker ComposeGitHub Actions

The clinic records are synthetic; the hospital labels come from the MIMIC-IV Clinical Database Demo (PhysioNet, ODbL v1.0). LOINC® is a registered trademark of Regenstrief Institute, Inc.