# From FHIR to OMOP CDM: building a research-grade health data pipeline

> An integration engineer’s walkthrough of a FHIR to OMOP pipeline: Bulk Data export, staging, vocabulary mapping with Athena, loading the CDM tables, and data quality checks, with the pitfalls that break studies.

Author: Ala Ben Aicha

Canonical page: https://alabenaicha.me/insights/fhir-to-omop-pipeline

Updated: 2026-10-06

## Direct answer

A FHIR to OMOP pipeline pulls FHIR resources in bulk, lands them unchanged in a staging area, maps every code to an OHDSI standard concept, and loads the result into OMOP CDM tables chosen by the concept's domain, not by the FHIR resource type. Patient becomes PERSON, Encounter becomes VISIT\_OCCURRENCE, Condition becomes CONDITION\_OCCURRENCE, and Observation is split between MEASUREMENT and OBSERVATION. Target OMOP CDM v5.4 unless your research network has moved to v5.5, keep the original codes in the `_source_value` and `_source_concept_id` fields, and run the Data Quality Dashboard before anyone runs a study.

## Why OMOP, and which version

FHIR is built for exchange: one patient, one transaction, current state. OMOP is built for analysis across millions of patients with identical SQL at every site. That is why EMA's [DARWIN EU](https://www.ema.europa.eu/en/about-us/how-we-work/big-data/data-analysis-real-world-interrogation-network-darwin-eu) asks its data partners to standardise into the OMOP common data model, and why a hospital that already exposes FHIR APIs often wants an OMOP copy for research.

Version choice in October 2026:

* The [OMOP CDM site](https://ohdsi.github.io/CommonDataModel/) lists **v5.5** as the current version. Its [changes from v5.4](https://ohdsi.github.io/CommonDataModel/cdm55Changes.html) are additive: seven optional fields (for example `value_as_source_concept_id` on MEASUREMENT), three vocabulary metadata tables, nothing removed or renamed.
* The HL7 [Vulcan FHIR to OMOP Implementation Guide](https://hl7.org/fhir/uv/omop/INFORMATIVE1/) (v1.0.0, informative, published by HL7's Biomedical Research and Regulation work group) targets **v5.4**, and its [CDM background page](https://hl7.org/fhir/uv/omop/INFORMATIVE1/en/the-omop-cdm.html) notes that the OHDSI community stopped developing v6.0 in 2022.

In practice: build to v5.4 columns, which a v5.5 database also accepts, and add the v5.5 fields when your network and tools expect them.

## Pipeline architecture

| Stage             | What happens                                                 | Typical tooling                                     |
| ----------------- | ------------------------------------------------------------ | --------------------------------------------------- |
| 1. Extract        | Group-level Bulk Data export, incremental with `_since`      | FHIR server `$export`, SMART Backend Services token |
| 2. Stage          | NDJSON landed as-is, then flattened to staging tables        | Object storage, Parquet, SQL views                  |
| 3. Map vocabulary | Source codes to standard concepts; local codes reviewed      | Athena vocabularies, Usagi, custom concepts         |
| 4. Load           | Domain routing, integer keys, OBSERVATION\_PERIOD derivation | SQL or dbt, Spark for volume                        |
| 5. Check          | Conformance, completeness, plausibility                      | Data Quality Dashboard, ACHILLES                    |

Extraction uses the [FHIR Bulk Data Access](https://hl7.org/fhir/uv/bulkdata/) kick-off. Exporting a research cohort Group keeps you away from whole-system dumps:

```http
GET /fhir/Group/research-cohort-01/$export?_type=Patient,Encounter,Condition,Observation,MedicationRequest,Procedure&_since=2026-09-01T00:00:00Z HTTP/1.1
Host: fhir.example.org
Accept: application/fhir+json
Prefer: respond-async
Authorization: Bearer <backend-services-token>
```

The server answers `202 Accepted` with a `Content-Location` status URL; you poll it until it returns a manifest of NDJSON files per resource type. Server support varies, so check before you design around it (see [FHIR servers compared](https://alabenaicha.me/insights/fhir-servers-compared)).

Keep the raw NDJSON. When a mapping rule changes six months later, you will rerun stages 3 to 5 from staging, not re-extract from production.

## Resource-to-table mapping

| FHIR resource                           | OMOP table                          | Key target fields                                                      | Watch out for                                                  |
| --------------------------------------- | ----------------------------------- | ---------------------------------------------------------------------- | -------------------------------------------------------------- |
| Patient                                 | PERSON (+ DEATH)                    | `gender_concept_id`, `year_of_birth`, `person_source_value`            | Never put the real MRN or national ID in `person_source_value` |
| Encounter                               | VISIT\_OCCURRENCE (+ VISIT\_DETAIL) | `visit_concept_id` from `Encounter.class`, start and end dates         | Ward transfers belong in VISIT\_DETAIL                         |
| Condition                               | CONDITION\_OCCURRENCE               | `condition_concept_id`, start date from onset or recorded date         | Filter `entered-in-error` and refuted verification status      |
| Observation                             | MEASUREMENT or OBSERVATION          | `measurement_concept_id`, `value_as_number`, `unit_concept_id`         | Routing depends on the standard concept's domain               |
| MedicationRequest / MedicationStatement | DRUG\_EXPOSURE                      | `drug_concept_id` (RxNorm or RxNorm Extension), `drug_type_concept_id` | A prescription is not proof the drug was taken                 |
| Procedure                               | PROCEDURE\_OCCURRENCE               | `procedure_concept_id`, `procedure_date`                               | Exclude planned or not-done procedures                         |

The IG also maps Immunization to DRUG\_EXPOSURE and AllergyIntolerance to OBSERVATION. Two tables have no FHIR resource behind them and still matter: OBSERVATION\_PERIOD (required by most OHDSI analytics, so derive it from first and last recorded activity, or from enrolment dates if you have them) and CDM\_SOURCE.

## Concept mapping: `source_concept_id` versus `concept_id`

Every clinical table carries two concept columns, and mixing them up is the most common ETL bug.

* `measurement_source_concept_id` holds the concept for the code you received, standard or not. A LOINC code goes here, and so does a local lab code that you registered as a custom concept.
* `measurement_concept_id` holds the **standard** concept reached by following the `Maps to` relationship in CONCEPT\_RELATIONSHIP. Analytics run on this column. If there is no mapping, it is `0`, never NULL and never the source concept.

The standard concept's `domain_id` decides the table. A SNOMED CT code sent in a Condition resource can map to an Observation-domain concept (a "history of" finding, for example) and must land in OBSERVATION, not CONDITION\_OCCURRENCE. Upstream code quality decides how much of this is automatic; the mapping layer described in [HL7 v2 to FHIR with LOINC and SNOMED CT](https://alabenaicha.me/insights/hl7v2-fhir-loinc-snomed-mapping) pays off twice here.

## Worked example: one lab result into MEASUREMENT

A synthetic FHIR Observation, already pseudonymised in staging:

```json
{
  "resourceType": "Observation",
  "id": "obs-000123",
  "status": "final",
  "code": { "coding": [{ "system": "http://loinc.org", "code": "2345-7", "display": "Glucose [Mass/volume] in Serum or Plasma" }] },
  "subject": { "reference": "Patient/pseudo-7f3a" },
  "encounter": { "reference": "Encounter/pseudo-enc-91" },
  "effectiveDateTime": "2026-09-14T08:30:00+02:00",
  "valueQuantity": { "value": 104, "unit": "mg/dL", "system": "http://unitsofmeasure.org", "code": "mg/dL" }
}
```

After flattening into `staging.observation`, the load resolves concepts from the vocabulary tables instead of hard-coding IDs (PostgreSQL syntax):

```sql
-- Synthetic example: staged FHIR Observations into OMOP CDM v5.4 MEASUREMENT
INSERT INTO cdm.measurement (
  measurement_id, person_id, measurement_concept_id, measurement_date,
  measurement_datetime, measurement_type_concept_id, value_as_number,
  unit_concept_id, visit_occurrence_id, measurement_source_value,
  measurement_source_concept_id, unit_source_value
)
SELECT
  nextval('cdm.measurement_id_seq'),
  p.person_id,
  COALESCE(std.concept_id, 0),
  CAST(o.effective_at AS date),
  o.effective_at,
  32817,                                  -- Type Concept 'EHR'
  o.value_number,
  COALESCE(u.concept_id, 0),
  v.visit_occurrence_id,
  o.code,                                 -- original source code
  COALESCE(src.concept_id, 0),
  o.unit_code
FROM staging.observation o
JOIN cdm.person p
  ON p.person_source_value = o.patient_pseudo_id
LEFT JOIN cdm.visit_occurrence v
  ON v.visit_source_value = o.encounter_pseudo_id
LEFT JOIN vocab.concept src
  ON src.vocabulary_id = 'LOINC' AND src.concept_code = o.code
LEFT JOIN vocab.concept_relationship cr
  ON cr.concept_id_1 = src.concept_id
 AND cr.relationship_id = 'Maps to'
 AND cr.invalid_reason IS NULL
LEFT JOIN vocab.concept std
  ON std.concept_id = cr.concept_id_2 AND std.standard_concept = 'S'
LEFT JOIN vocab.concept u
  ON u.vocabulary_id = 'UCUM' AND u.concept_code = o.unit_code
WHERE o.code_system = 'http://loinc.org'
  AND o.status IN ('final', 'amended', 'corrected')
  AND COALESCE(std.domain_id, src.domain_id) = 'Measurement';
```

Rows that fail the domain filter go to the OBSERVATION load; codes found in neither join go to a mapping backlog, not silently into concept `0` forever. The `32817` value is the EHR type concept the IG's own measurement StructureMap uses; confirm it against your Athena release.

## OHDSI tooling you will actually use

| Tool                                | Role in the pipeline                                                                                    |
| ----------------------------------- | ------------------------------------------------------------------------------------------------------- |
| [Athena](https://athena.ohdsi.org/) | Download the standardised vocabularies (SNOMED CT, LOINC, RxNorm, UCUM and others; some need a licence) |
| WhiteRabbit                         | Profile source tables before mapping                                                                    |
| Rabbit-in-a-Hat                     | Document table and field mappings; it produces specs, not code                                          |
| Usagi                               | Suggest mappings for local codes by text similarity; a human approves                                   |
| ACHILLES                            | Characterise the loaded CDM for review and for ATLAS                                                    |
| Data Quality Dashboard              | Run conformance, completeness and plausibility checks table by table                                    |

All are listed on the [OHDSI software page](https://www.ohdsi.org/software-tools/). Treat a DQD run as a release gate, not a report someone reads later.

## Pitfalls that break studies

* **Local codes.** European lab and drug codes are often local. Register them as [custom concepts](https://ohdsi.github.io/CommonDataModel/customConcepts.html) (IDs above 2,000,000,000, never standard, used only in `_source_concept_id` fields) and map them to standard concepts with `Maps to`. Unmapped local codes are invisible to network studies.
* **Units.** Use the UCUM code from `valueQuantity.code`, not the display `unit`. UCUM is case-sensitive, and `mmol/l` from a legacy feed will not match. Do not convert values silently; if you normalise, document the rule.
* **Status and intent.** FHIR carries drafts, plans, cancellations and errors. Filter them deliberately. The IG's [common challenges page](https://hl7.org/fhir/uv/omop/INFORMATIVE1/en/F2OGeneralIssues.html) covers status, intent, identifiers and temporal precision.
* **Date shifting.** A consistent per-person offset preserves intervals but breaks seasonality, calendar-based exposure windows and anything tied to a real date such as a vaccine campaign. Agree on it with the study team before you load.
* **Pseudonymisation under GDPR.** Pseudonymised data is still personal data under the [GDPR](https://eur-lex.europa.eu/eli/reg/2016/679/oj). Keep the re-identification key outside the OMOP environment, scrub free-text `_source_value` fields, and read the OHDSI [privacy guidance](https://ohdsi.github.io/CommonDataModel/cdmPrivacy.html). The architecture side is in [GDPR-compliant healthcare data architecture](https://alabenaicha.me/insights/gdpr-compliant-healthcare-data-architecture).
* **Unstable keys.** OMOP keys are integers. Keep a persistent crosswalk from FHIR logical IDs to OMOP IDs so incremental loads update rather than duplicate.

## The European angle

The [EHDS Regulation](https://health.ec.europa.eu/ehealth-digital-health-and-care/european-health-data-space-regulation-ehds_en) entered into force on 26 March 2025, and its secondary-use rules apply from March 2029 for most data categories, with access granted through health data access bodies. The regulation does not mandate OMOP. But a data holder that can already produce a documented, quality-checked OMOP dataset from its FHIR feeds will answer access requests faster, and can join DARWIN EU-style federated studies without a separate project. The EHDS context is in the [EHDS guide](https://alabenaicha.me/insights/european-health-data-space-ehds-guide).

If you are planning a FHIR to OMOP pipeline and want help with the extraction, vocabulary mapping or quality gates, that is the scope of [healthcare data analytics work](https://alabenaicha.me/services/healthcare-data-analytics).
