01Why it exists
A fasting glucose result is GLUC at one lab, FASTING BLOOD SUGAR, SER/PLAS
at the next and L-002 at a third. Each is correct inside its own building and
unreadable outside it. The result cannot follow the patient to another institution,
national reporting counts one measurement three different ways, and a test that already
happened gets ordered again because nobody downstream can tell. FHIR can carry that result
anywhere, but only once it is coded as LOINC.
Ontora does that coding at scale. Records arrive over REST or Kafka. An embedding model
proposes LOINC candidates, and every candidate is checked against a real FHIR terminology
server. A reviewer decides, Ontora remembers the decision for that clinic, and publishes
it as a standard FHIR ConceptMap.
The mapping is not the valuable part; a lookup table can map. What matters is that the system has been measured against hidden ground truth, so it knows which of its own answers to trust.
02A confidence score you can check
The system was run over 660 evaluation records: 600 from three simulated clinics and 60 genuine de-identified hospital labs from MIMIC-IV. Each carries a ground-truth LOINC code that is never exposed through the API and never shown to the model. The 655 that have one are scored; the other five belong to lab items the 91-concept subset cannot honestly code.
Top-1 precision per confidence band, scored against hidden ground truth.
A cosine similarity is not a probability
The calibration chart shows it plainly. Expected calibration error is 17.6% over equal-count deciles, or 18.2% with equal-width bins. The shipped HIGH cut-off of 0.85 is deliberately conservative: the lowest cut-off that still holds 95% precision is 0.772, and the cost of every setting is tabulated by a script rather than typed by hand.
03Where it fails, on purpose
The same numbers, split by what each source sends:
That last row is the argument for the confidence spread. L-002 carries no
meaning, so no model can recover it, and Ontora does not pretend otherwise: it scores those
records LOW and sends them to a person. Once a person decides what L-002 means
at St. Mary’s, the mapping memory applies that decision to every later
L-002 from St. Mary’s without asking the model again, and the result still
names the reviewer who made the original call as its author.
04Architecture
ingest · review · audit · FHIR
| Service | Why it is its own thing |
|---|---|
| ontora-core | The only component that writes. REST and Kafka are two adapters over one service, so both paths are audited identically. |
| ontora-matcher | The only part that needs PyTorch. Stateless, with 468 name vectors for 91 concepts in memory, so it scales and redeploys on its own schedule. |
| HAPI FHIR JPA | A real terminology server, so validation is an external check rather than a list the system grades itself against. |
| PostgreSQL 16 | Five tables, FHIR resources as jsonb, eleven Flyway migrations. Hibernate only validates the schema; it never changes it. |
| Kafka | A lab feed is a stream, not a request. Partitioning by clinic keeps each clinic’s results in order, and matching consumer concurrency to partitions is where 42 → 68 rec/s came from. |
One record’s journey
- A clinic sends GLUC · 5.4 mmol/L · Patient/0042, over REST or Kafka.
- The core asks whether this clinic’s GLUC has been decided before.
- If it has, the mapping memory applies 1558-6 straight away and skips to the last step.
- Otherwise the matcher returns three candidates with their similarity.
- Every candidate is checked with $validate-code against HAPI FHIR, then gets a confidence band, a specimen check and a UCUM unit check.
- The reviewer’s queue updates live over server-sent events.
- The reviewer approves 1558-6; it is validated again at approval and remembered for this clinic.
- A FHIR Observation, Provenance and AuditEvent are stored. The reviewer’s identity comes from the JWT, never from the request body.
05The product
Amber means not yet portable, green means portable. Nothing on these screens describes a patient; they describe whether a result can be read somewhere else.
06Engineering decisions
Each of these had a real alternative.
-
01 · MODEL CHOICE
MiniLM over SapBERT
SapBERT ranks 1.7 points better on top-1, but it is about five times slower, adds 420 MB, and makes the HIGH band less precise. That band carries the safety argument, so MiniLM stays the default.
-
02 · TRANSACTIONS
Publishing is its own action
Publishing the
ConceptMapis not part of approval. Otherwise a network call to HAPI sits inside the approval transaction, and approvals fail whenever the terminology server is down. The cost is one manual step. -
03 · CHECKS
Warn, never block
A candidate with a unit warning is right 2.5% of the time, but the check is a heuristic over free text, and a heuristic that blocks loses records silently. Flagged candidates are kept out of bulk approval, so a person always sees them.
-
04 · SIGNALS
The matcher never sees the axes
It embeds names only; specimen and UCUM checks run afterwards in Kotlin. Two independent signals are useful precisely when they disagree, and folding them into one score would hide that.
-
05 · TENANCY
Memory is per clinic
GLUCat one lab is a statement about that lab. Mappings are unique per clinic and source code, published under per-clinic code systems, and never shared across clinics. -
06 · SECURITY
Duties split by the server
Admin and reviewer roles never overlap: an admin gets a 403 on approve, a reviewer a 403 on the audit trail. Both are gated in the security filter chain, because method security throws where the catch-all handler would turn a 403 into a 500.
07Tests
JUnit 5 with Testcontainers and REST Assured for the core, against real PostgreSQL and
Kafka containers; Vitest for the Angular UI; pytest with the real model for the matcher.
Every FHIR resource Ontora produces (Observation, Provenance,
AuditEvent, ConceptMap and the rest) is validated offline against
the official R4 profiles, so a malformed resource fails the build instead of reaching an
exchange partner. GitHub Actions runs the three suites as independent jobs.
08Stated limits, and what’s next
- The throughput figures were taken on a single laptop with every service running on it.
- Accounts are held in memory, where a deployment would use an identity provider.
- The live-update stream takes its token in the query string, because
EventSourcecannot set headers. - One clinic key is shared by every feed, where a deployment would issue one per feed.
What separates this from something that ships, in the order they unblock the most:
The clinic records are synthetic; the hospital labels come from the MIMIC-IV Clinical Database Demo (PhysioNet, ODbL v1.0). LOINC® is a registered trademark of Regenstrief Institute, Inc.


