All documentation Download (Markdown) Technical

CDR 5 · cdr5.pdhc.se

Technical manual


cdr.pdhc — Technical Manual

One codebase, five instances (CDR 1–5). All five run the identical
software and schema. They differ only in configuration — chiefly the
CDR_READ_LOCKDOWN flag (see §3) — and in the data they hold.

1. Overview

cdr.pdhc is the platform's clinical data repository. It ingests normalised
observations from the gateways, stores every FHIR resource type in its own
per-type table (with full version history), exposes a FHIR R5 read/search
surface plus a care-delivery clinical surface, and — on CDR 1 only —
forwards concept-mapped data to the Cambio CDR sandbox.

Ports: 9046 (Flask/gunicorn, bound to 127.0.0.1), 9047 (PostgreSQL),
9048–9049 reserved. Health at GET /healthz (also GET /api/v1/health).

2. Storage model — per-FHIR-type tables

The authoritative store is a set of per-resource-type tables, one live
table per FHIR type plus a matching *_history twin. This replaced the old
three-layer (raw → standard → canonical) design; the old tables still
exist but are legacy (see §2.3).

2.1 Per-type live + history tables

app/models/resources.py builds, from a single RESOURCES list, one live
ORM model and one history model for each of these ten FHIR R5 types:

FHIR type Live table History table
Patient patient patient_history
Observation observation observation_history
QuestionnaireResponse questionnaire_response questionnaire_response_history
Condition condition condition_history
MedicationStatement medication_statement medication_statement_history
MedicationRequest medication_request medication_request_history
AllergyIntolerance allergy_intolerance allergy_intolerance_history
Procedure procedure procedure_history
Encounter encounter encounter_history
DiagnosticReport diagnostic_report diagnostic_report_history

Every live table shares these common columns:

guid, patient_guid, org_guid, code_canonical, effective_at, raw_json,
source, source_request_id, meta_tag, version_id, sync_group_id,
mapping_version, etag, received_at, created_at, updated_at

plus a few type-specific columns (e.g. observation adds value_quantity,
value_unit, value_string, value_code, status; patient adds
identifiers, names, gender, birth_date, active). The full FHIR
resource is always kept verbatim in raw_json; the extracted columns exist
only to index and search.

Time fields (#294 RFC E1): effective_at is the clinical measurement
time; received_at is when the platform first saw the payload; created_at
is when this row was written; updated_at is the last version bump.

History tables carry the same columns (minus the live-only ones) plus
superseded_at and superseded_by_request_id. The primary key is the
composite (guid, version_id) — one row per historical version.

code_canonical is the concept key a row is searched by. On live CDR 1
it is dominantly the Path-B embedded-GUID form urn:pdhc:concept/<guid>
— the plan.pdhc Concept GUID is the last path segment. Display resolution
parses that GUID out and looks it up via plan.pdhc (see §6.2).

2.2 Cross-cutting tables

2.3 Legacy tables (read-mostly)

The pre-platform-plan tables remain in place: ingest_raw, fhir_resources,
openehr_compositions, health_observations, activities, plus
clinical_context, dedupe_registry, loinc_archetype_map, service_keys,
users, audit_log, cambio_patient_map, cambio_delivery_log.

The flat ingest path (§4.1) still writes several of these — ingest_raw
(immutable payload + SHA-256), fhir_resources, an openehr_compositions
row generated by the transformer, clinical_context, dedupe_registry,
cambio_delivery_log, audit_log — because they back dedup, provenance and
Cambio delivery. But the read/search surface no longer reads them: step
3.5 of the flat ingest (#295) mirrors each Observation into the per-type
observation table, and all FHIR reads hit the per-type tables. So the
legacy standard/canonical tables (fhir_resources, openehr_compositions,
health_observations, activities) are effectively write-through legacy for
provenance, not the query store.

Table count is ~37, not 13: 10 live + 10 history + 4 cross-cutting
(sync_group, change_feed, cdr_audit_plan_miss, cdr_read_audit) + 13
legacy/infra tables.

3. The five CDRs and CDR_READ_LOCKDOWN (#293)

The single most important config difference between instances:

Under lockdown (_is_read_path in app/auth.py), the read endpoints
(/api/v1/fhir/*, /api/v1/openehr/*, /api/v1/stats, /api/v1/cambio/*,
/api/v1/clinical/*) accept only the analyse-layer reader identities
dashboard.pdhc and analyse.pdhc. All other trusted services may still
write; they just cannot read a locked-down CDR. CDR 1, with the flag
false, serves reads to any valid trusted reader.

4. Data flow

4.1 Flat ingest (legacy wire, gateway → CDR 1)

POST /api/v1/ingest and POST /api/v1/ingest/batch (max 100). Service-key
auth via KNOWN_SERVICES (§5.1). Body must carry patient_guid and may
carry fhir_resource, openehr_composition, canonical, and
clinical_context blocks. IngestPipeline.process runs:

  1. Deduplicate — SHA-256 of the payload vs dedupe_registry (scoped by
    source service). A repeat returns 200 {"status":"duplicate"}.
  2. Store raw — immutable ingest_raw insert.
  3. Store standard — write the FHIR resource to fhir_resources; store or
    (from FHIR) generate an openehr_compositions row via the transformer
    (§6.1). If only openEHR arrived, generate the FHIR side.
    3.5 Mirror to per-type (#295) — build and insert the live
    observation row (this is what search reads).
  4. Store canonicalhealth_observations or activities (legacy).
  5. Store context — the canonical 12-field clinical_context row.
  6. Register dedupe — add the hash to dedupe_registry.
  7. Enqueue Cambio — create cambio_delivery_log rows (pending if a
    concept reference is present, else skipped). The X2 operator session id
    (X-Operator-Session-Id) is captured for later replay.
  8. Audit — write audit_log with correlation id and client IP.

Returns 202 accepted, 200 duplicate, or 422 rejected.

GET /api/v1/ingest/by-source-id/<source_system_id> (service-key,
source-scoped) confirms a prior delivery landed.

4.2 FHIR write (per-type path)

POST /api/v1/fhir/<Type>, PUT /api/v1/fhir/<Type>/<guid> (optimistic
concurrency via If-Match), and POST /api/v1/fhir/Bundle (transaction =
atomic; batch = per-entry). Each entry runs:

canonicalise → dedup-lookup → insert-or-update-with-history →
sync_group → mapping_version → change_feed

Canonicalisation resolves codings to a code_canonical (via xlate.pdhc /
plan.pdhc). A body carrying an id upserts by that id (FHIR "update as
create"), version-bumping and pushing the prior row to *_history. Error
outcomes: 422 xlate_miss, 422 plan_miss (also bookkept to
cdr_audit_plan_miss), 412 on stale If-Match, 503 transient when xlate
or plan is unreachable.

All under /api/v1/fhir/ (app/api/fhir_read.py), reading the per-type
tables:

Endpoint Purpose
GET /<Type> Search (see params below)
GET /<Type>/<guid> Instance read (ETag)
GET /<Type>/<guid>/_history Version list Bundle
GET /<Type>/<guid>/_history/<vid> vread a specific version
GET /Patient/<guid>/$everything Patient compartment Bundle
GET /events change_feed long-poll (?since=<seq>)
GET /metadata FHIR CapabilityStatement

Search parameters are FHIR-standard (not the old patient_guid /
loinc_code):

Every read is org-scoped (Rule 24, §5.3), consent-filtered (§5.4), and
written to cdr_read_audit.

Group aggregations ($stats, $agp) were moved out to the analyse layer
(dashboard.pdhc) in the CDR1/analyse split (#289). CDR 1 is pure storage;
analyse fetches raw Observations via search and aggregates locally.

4.4 Care-delivery clinical read surface (#468)

app/api/clinical_read.py, mounted at /api/v1/clinical. This serves the
rebuilt single-patient clinical dashboard (#462), which reads CDR 1 under
a care-delivery legal basis (vårdrelation + spärr), not the
analysis-consent basis:

Endpoint Returns
GET /clinical/patients The org's patients that have Observation data, with counts
GET /clinical/patient/<guid>/summary Per-concept counts + unit + first/last seen
GET /clinical/patient/<guid>/series Time-series points (optional code, from, to filters)

Guard (_care_delivery_guard): the request must carry
X-Access-Purpose: care-delivery and the service identity must be
dashboard.pdhc. Because the service blob is is_su_admin, the shared FHIR
org-filter would leak all orgs — so these endpoints do their own explicit
org scoping from the forwarded X-Org-Guids / X-Is-Admin headers (the
operator's affiliation care-unit guids). Consent (#422) is bypassed here:
a patient may be treatable while having declined research. Series points
carry org_guid so the dashboard can apply spärr (per-clinic blocks) on its
side. Concept display names are resolved via plan.pdhc (#471), fail-open.

4.5 openEHR read

GET /api/v1/openehr/composition/<guid> — a single by-GUID read of a legacy
openehr_compositions row. The former search form
(?patient_guid=&archetype_id=) was moved to the analyse layer (#292); CDR 1
keeps only the storage-style per-GUID lookup.

4.6 Other read endpoints

5. Authentication and access control

5.1 Ingest (write) service keys — KNOWN_SERVICES

Header pair X-Source-Service + X-Service-Key (app/api/auth.py).
Accepted sources and the env var each key is matched against:

Source service Config / env var
gateway.pdhc GATEWAY_PDHC_SERVICE_KEY
2gate.pdhc TWOGATE_PDHC_SERVICE_KEY
sim.pdhc SIM_PDHC_SERVICE_KEY

5.2 FHIR/read trusted services — KNOWN_FHIR_SERVICES

Sibling services may read/write the FHIR surface with a service key
(app/auth.py), each keyed to its own env var:

Source service Config / env var Role
sim.pdhc SIM_PDHC_SERVICE_KEY writes synthetic cohorts
dashboard.pdhc DASHBOARD_PDHC_SERVICE_KEY analyse-layer + clinical reader
analyse.pdhc ANALYSE_PDHC_SERVICE_KEY extracted analyse-layer reader (#541)

Under CDR_READ_LOCKDOWN, only dashboard.pdhc and analyse.pdhc may take
a read path; every listed service may still write. A valid
service-key request gets a synthetic is_su_admin blob tagged with
service_source, which downstream code uses to distinguish machine writes
from human operators.

5.3 SSO (human operators)

AUTH_MODE=sso validates the session bearer token against
sso.pdhc /api/auth/me/service on every request (no blob caching, so an
SSO logout takes effect immediately). Access requires the analysis phase gate
is_su_admin, or user_type == "professional" with analysis in the
session phases. flask create-su bootstraps an SU. AUTH_MODE=off (local
dev only) loads a dev SU blob.

Org scoping (Rule 24): non-admin reads are filtered to the operator's
Zone-1 care-unit guids — affiliations[].care_unit_guid, with a dual-read
fallback to the legacy organization_ids (M0 #416). Admins (is_su_admin)
see all.

app/services/analysis_consent.py enforces analysis-phase consent (EHDS
opt-out, per-project research consent, quality-registry opt-out). The verdict
comes from ips.pdhc (POST /api/v1/patients/analysis-filter) — nothing
is computed locally. The operator's purpose is derived from their active
affiliation role
(researcher → research; quality/registry → quality_registry;
other clinical → statistics; SU-admin without affiliations → administration,
never blocked). check_patient_allowed gates a single patient (403 on
exclusion); consent_allowed_guids batch-filters a search result set. Fail
closed:
if ips is unreachable, reads abort 503. Service-key/machine
contexts and the care-delivery surface (§4.4) pass through — the sibling
holds the real operator context.

6. FHIR ↔ openEHR transformation

6.1 Flat-ingest transformer

app/services/transformer.py generates an openEHR composition from a FHIR
Observation on the flat-ingest path, using a LOINC-to-archetype seed map (11
entries: body weight, blood pressure, pulse, temperature, SpO2, height,
respiration, BMI, blood glucose, waist circumference, sleep). Unknown LOINC
codes fall back to openEHR-EHR-OBSERVATION.laboratory_test_result.v1. This
is the only path that materialises openEHR; the per-type write path defers
openEHR and only mints the sync_group_id.

6.2 Concept display (#471)

clinical_read.resolve_display parses the Concept GUID out of a Path-B
code_canonical (urn:pdhc:concept/<guid>) and resolves a human label via
plan.pdhc CodeSystem/$lookup (cached in PlanClient). Cosmetic and
fail-open — a miss shows the raw code.

7. Cambio CDR sandbox delivery (CDR 1)

Only observations with a concept reference are eligible; unmapped data stays
local (cambio_delivery_log.status = skipped). When
CAMBIO_DELIVERY_ENABLED=true, an APScheduler background worker runs a
delivery cycle every 60 s: ensure the patient exists in Cambio (FHIR Patient
+ openEHR EHR), deliver the FHIR Observation and openEHR Composition, retry
with backoff. The X2 operator session id captured at ingest is replayed as
X-Operator-Session-Id on the CDR 1 → Cambio hop so chain-of-custody
survives the async gap. Status/retry via /api/v1/cambio/*.

8. Health check

GET /healthz (and GET /api/v1/health) returns:

{ "status": "ok", "service": "cdr.pdhc", "database": "connected" }

status is ok (HTTP 200) when the DB SELECT 1 succeeds, degraded
(HTTP 503) otherwise; database is connected / unavailable. There is
no version field.
CORS headers allow https://www.pdhc.se/services.html
to read the body cross-origin and drive the real status/DB dots.

9. Configuration (environment)

Variable Purpose
DATABASE_URL PostgreSQL connection string (default port 9047)
AUTH_MODE off (dev) or sso (production)
SSO_BASE_URL / SSO_CLIENT_ID / SSO_CLIENT_SECRET / SSO_CALLBACK_URL SSO validation
CDR_READ_LOCKDOWN false on CDR 1, true on CDR 2–5 (#293)
GATEWAY_PDHC_SERVICE_KEY / TWOGATE_PDHC_SERVICE_KEY / SIM_PDHC_SERVICE_KEY Ingest service keys
DASHBOARD_PDHC_SERVICE_KEY / ANALYSE_PDHC_SERVICE_KEY Read/analyse service keys
PLAN_BASE_URL plan.pdhc — canonicalisation + concept display
IPS_BASE_URL ips.pdhc — consent analysis-filter (#422)
STRICT_CANONICALISATION Relax plan-validate on transient unreachability when false
CAMBIO_DELIVERY_ENABLED Activate the delivery worker (CDR 1)
CAMBIO_CLIENT_ID / CAMBIO_CLIENT_SECRET / CAMBIO_TOKEN_URL / CAMBIO_BASE_URL / CAMBIO_TENANT / CAMBIO_*_HSA_ID Cambio sandbox credentials

10. Operations

Cold start bash start.sh. Graceful restart of the CDR's own service only
(Rule 19 / CLAUDE.md §15) — never touch sibling services, volumes, or the
shared Colima VM. See the platform CLAUDE.md for the deploy layout, the
Postgres credential-drift trap, and the Docker/Colima gotchas.

Port Allocation

All ports bind to 127.0.0.1 (loopback only); external traffic arrives
via the reverse proxy. Every instance runs identical software; only the
host port pair (APP_PORT / DB_PORT) differs. Inside every container
the app listens on 9046 and Postgres on 5432.

Instance App (Gunicorn) PostgreSQL
CDR 1 9046 9045 (local-dev default 9047)
CDR 2 9146 9145
CDR 3 9246 9245
CDR 4 9346 9345
CDR 5 9446 9445