Healthcare AI-7 min read

From FHIR to OMOP CDM: building a research-grade health data pipeline

An integration engineer’s walkthrough of a FHIR to OMOP pipeline: Bulk Data export, staging, vocabulary mapping with Athena, loading the CDM tables, and data quality checks, with the pitfalls that break studies.

Ala Ben Aicha

From FHIR to OMOP CDM: building a research-grade health data pipeline

Direct answer

A FHIR to OMOP pipeline pulls FHIR resources in bulk, lands them unchanged in a staging area, maps every code to an OHDSI standard concept, and loads the result into OMOP CDM tables chosen by the concept's domain, not by the FHIR resource type. Patient becomes PERSON, Encounter becomes VISIT_OCCURRENCE, Condition becomes CONDITION_OCCURRENCE, and Observation is split between MEASUREMENT and OBSERVATION. Target OMOP CDM v5.4 unless your research network has moved to v5.5, keep the original codes in the _source_value and _source_concept_id fields, and run the Data Quality Dashboard before anyone runs a study.

Why OMOP, and which version

FHIR is built for exchange: one patient, one transaction, current state. OMOP is built for analysis across millions of patients with identical SQL at every site. That is why EMA's DARWIN EU asks its data partners to standardise into the OMOP common data model, and why a hospital that already exposes FHIR APIs often wants an OMOP copy for research.

Version choice in October 2026:

  • The OMOP CDM site lists v5.5 as the current version. Its changes from v5.4 are additive: seven optional fields (for example value_as_source_concept_id on MEASUREMENT), three vocabulary metadata tables, nothing removed or renamed.
  • The HL7 Vulcan FHIR to OMOP Implementation Guide (v1.0.0, informative, published by HL7's Biomedical Research and Regulation work group) targets v5.4, and its CDM background page notes that the OHDSI community stopped developing v6.0 in 2022.

In practice: build to v5.4 columns, which a v5.5 database also accepts, and add the v5.5 fields when your network and tools expect them.

Pipeline architecture

Stage What happens Typical tooling
1. Extract Group-level Bulk Data export, incremental with _since FHIR server $export, SMART Backend Services token
2. Stage NDJSON landed as-is, then flattened to staging tables Object storage, Parquet, SQL views
3. Map vocabulary Source codes to standard concepts; local codes reviewed Athena vocabularies, Usagi, custom concepts
4. Load Domain routing, integer keys, OBSERVATION_PERIOD derivation SQL or dbt, Spark for volume
5. Check Conformance, completeness, plausibility Data Quality Dashboard, ACHILLES

Extraction uses the FHIR Bulk Data Access kick-off. Exporting a research cohort Group keeps you away from whole-system dumps:

GET /fhir/Group/research-cohort-01/$export?_type=Patient,Encounter,Condition,Observation,MedicationRequest,Procedure&_since=2026-09-01T00:00:00Z HTTP/1.1
Host: fhir.example.org
Accept: application/fhir+json
Prefer: respond-async
Authorization: Bearer <backend-services-token>

The server answers 202 Accepted with a Content-Location status URL; you poll it until it returns a manifest of NDJSON files per resource type. Server support varies, so check before you design around it (see FHIR servers compared).

Keep the raw NDJSON. When a mapping rule changes six months later, you will rerun stages 3 to 5 from staging, not re-extract from production.

Resource-to-table mapping

FHIR resource OMOP table Key target fields Watch out for
Patient PERSON (+ DEATH) gender_concept_id, year_of_birth, person_source_value Never put the real MRN or national ID in person_source_value
Encounter VISIT_OCCURRENCE (+ VISIT_DETAIL) visit_concept_id from Encounter.class, start and end dates Ward transfers belong in VISIT_DETAIL
Condition CONDITION_OCCURRENCE condition_concept_id, start date from onset or recorded date Filter entered-in-error and refuted verification status
Observation MEASUREMENT or OBSERVATION measurement_concept_id, value_as_number, unit_concept_id Routing depends on the standard concept's domain
MedicationRequest / MedicationStatement DRUG_EXPOSURE drug_concept_id (RxNorm or RxNorm Extension), drug_type_concept_id A prescription is not proof the drug was taken
Procedure PROCEDURE_OCCURRENCE procedure_concept_id, procedure_date Exclude planned or not-done procedures

The IG also maps Immunization to DRUG_EXPOSURE and AllergyIntolerance to OBSERVATION. Two tables have no FHIR resource behind them and still matter: OBSERVATION_PERIOD (required by most OHDSI analytics, so derive it from first and last recorded activity, or from enrolment dates if you have them) and CDM_SOURCE.

Concept mapping: source_concept_id versus concept_id

Every clinical table carries two concept columns, and mixing them up is the most common ETL bug.

  • measurement_source_concept_id holds the concept for the code you received, standard or not. A LOINC code goes here, and so does a local lab code that you registered as a custom concept.
  • measurement_concept_id holds the standard concept reached by following the Maps to relationship in CONCEPT_RELATIONSHIP. Analytics run on this column. If there is no mapping, it is 0, never NULL and never the source concept.

The standard concept's domain_id decides the table. A SNOMED CT code sent in a Condition resource can map to an Observation-domain concept (a "history of" finding, for example) and must land in OBSERVATION, not CONDITION_OCCURRENCE. Upstream code quality decides how much of this is automatic; the mapping layer described in HL7 v2 to FHIR with LOINC and SNOMED CT pays off twice here.

Worked example: one lab result into MEASUREMENT

A synthetic FHIR Observation, already pseudonymised in staging:

{
  "resourceType": "Observation",
  "id": "obs-000123",
  "status": "final",
  "code": { "coding": [{ "system": "http://loinc.org", "code": "2345-7", "display": "Glucose [Mass/volume] in Serum or Plasma" }] },
  "subject": { "reference": "Patient/pseudo-7f3a" },
  "encounter": { "reference": "Encounter/pseudo-enc-91" },
  "effectiveDateTime": "2026-09-14T08:30:00+02:00",
  "valueQuantity": { "value": 104, "unit": "mg/dL", "system": "http://unitsofmeasure.org", "code": "mg/dL" }
}

After flattening into staging.observation, the load resolves concepts from the vocabulary tables instead of hard-coding IDs (PostgreSQL syntax):

-- Synthetic example: staged FHIR Observations into OMOP CDM v5.4 MEASUREMENT
INSERT INTO cdm.measurement (
  measurement_id, person_id, measurement_concept_id, measurement_date,
  measurement_datetime, measurement_type_concept_id, value_as_number,
  unit_concept_id, visit_occurrence_id, measurement_source_value,
  measurement_source_concept_id, unit_source_value
)
SELECT
  nextval('cdm.measurement_id_seq'),
  p.person_id,
  COALESCE(std.concept_id, 0),
  CAST(o.effective_at AS date),
  o.effective_at,
  32817,                                  -- Type Concept 'EHR'
  o.value_number,
  COALESCE(u.concept_id, 0),
  v.visit_occurrence_id,
  o.code,                                 -- original source code
  COALESCE(src.concept_id, 0),
  o.unit_code
FROM staging.observation o
JOIN cdm.person p
  ON p.person_source_value = o.patient_pseudo_id
LEFT JOIN cdm.visit_occurrence v
  ON v.visit_source_value = o.encounter_pseudo_id
LEFT JOIN vocab.concept src
  ON src.vocabulary_id = 'LOINC' AND src.concept_code = o.code
LEFT JOIN vocab.concept_relationship cr
  ON cr.concept_id_1 = src.concept_id
 AND cr.relationship_id = 'Maps to'
 AND cr.invalid_reason IS NULL
LEFT JOIN vocab.concept std
  ON std.concept_id = cr.concept_id_2 AND std.standard_concept = 'S'
LEFT JOIN vocab.concept u
  ON u.vocabulary_id = 'UCUM' AND u.concept_code = o.unit_code
WHERE o.code_system = 'http://loinc.org'
  AND o.status IN ('final', 'amended', 'corrected')
  AND COALESCE(std.domain_id, src.domain_id) = 'Measurement';

Rows that fail the domain filter go to the OBSERVATION load; codes found in neither join go to a mapping backlog, not silently into concept 0 forever. The 32817 value is the EHR type concept the IG's own measurement StructureMap uses; confirm it against your Athena release.

OHDSI tooling you will actually use

Tool Role in the pipeline
Athena Download the standardised vocabularies (SNOMED CT, LOINC, RxNorm, UCUM and others; some need a licence)
WhiteRabbit Profile source tables before mapping
Rabbit-in-a-Hat Document table and field mappings; it produces specs, not code
Usagi Suggest mappings for local codes by text similarity; a human approves
ACHILLES Characterise the loaded CDM for review and for ATLAS
Data Quality Dashboard Run conformance, completeness and plausibility checks table by table

All are listed on the OHDSI software page. Treat a DQD run as a release gate, not a report someone reads later.

Pitfalls that break studies

  • Local codes. European lab and drug codes are often local. Register them as custom concepts (IDs above 2,000,000,000, never standard, used only in _source_concept_id fields) and map them to standard concepts with Maps to. Unmapped local codes are invisible to network studies.
  • Units. Use the UCUM code from valueQuantity.code, not the display unit. UCUM is case-sensitive, and mmol/l from a legacy feed will not match. Do not convert values silently; if you normalise, document the rule.
  • Status and intent. FHIR carries drafts, plans, cancellations and errors. Filter them deliberately. The IG's common challenges page covers status, intent, identifiers and temporal precision.
  • Date shifting. A consistent per-person offset preserves intervals but breaks seasonality, calendar-based exposure windows and anything tied to a real date such as a vaccine campaign. Agree on it with the study team before you load.
  • Pseudonymisation under GDPR. Pseudonymised data is still personal data under the GDPR. Keep the re-identification key outside the OMOP environment, scrub free-text _source_value fields, and read the OHDSI privacy guidance. The architecture side is in GDPR-compliant healthcare data architecture.
  • Unstable keys. OMOP keys are integers. Keep a persistent crosswalk from FHIR logical IDs to OMOP IDs so incremental loads update rather than duplicate.

The European angle

The EHDS Regulation entered into force on 26 March 2025, and its secondary-use rules apply from March 2029 for most data categories, with access granted through health data access bodies. The regulation does not mandate OMOP. But a data holder that can already produce a documented, quality-checked OMOP dataset from its FHIR feeds will answer access requests faster, and can join DARWIN EU-style federated studies without a separate project. The EHDS context is in the EHDS guide.

If you are planning a FHIR to OMOP pipeline and want help with the extraction, vocabulary mapping or quality gates, that is the scope of healthcare data analytics work.

OMOP CDMFHIROHDSIBulk DataETLLOINCSNOMED CTReal-world evidenceEHDSGDPR

Related reading and services

Let's Continue the Conversation

Have questions about this topic? I'd love to hear from you.