| Stage | SEER stage mix (known stage), % | Share written into both modules, % | First run (veteran module unpatched), % | Analysis run (both patched), % |
|---|---|---|---|---|
| I | 20.0 | 20.02 | 51.5 | 21.8 |
| II | 8.4 | 8.41 | 5.8 | 7.4 |
| III | 19.8 | 19.82 | 11.6 | 19.2 |
| IV | 51.7 | 51.75 | 31.1 | 51.5 |
How realistic is synthetic oncology EHR data?
NSCLC survival from Synthea, harmonized to OMOP and benchmarked against SEER
1 Summary
This project built an end-to-end workflow (Synthea → OMOP CDM 5.4 → dbt → R) on a synthetic NSCLC cohort of 1,347 patients and benchmarked its survival against 172,582 real U.S. patients from SEER (160,258 with known stage). Unmodified Synthea staged every lung cancer as stage I, and even after patching that module a second, hidden path (the veteran module) still produced 51.5% stage I, so both had to be fixed before any comparison was possible; the stage mix now matches SEER (chi-square p = 0.50 for age 50 and over). Synthea ranks the stages correctly, and some single summaries look close to SEER, such as the stage III median (13.2 vs 14 months) and the stage I restricted mean survival time (42.9 vs 43.5 months), yet the curves do not match: nobody with stage II–IV disease is followed beyond 28 months, whereas SEER five-year survival is 33%, 17% and 4% for stages II, III and IV. Age and sex have no effect on survival in Synthea (hazard ratio 1.05 per 10 years and 1.01 for men), reweighting to SEER’s age and sex distribution leaves the curves essentially unchanged, and 1,346 of 1,347 patients received the same treatment, so treatment effects cannot be identified. Synthetic data can test a pipeline but not stand in for real-world outcomes, and the lesson is to validate full distributions, not summary statistics.
2 Background and question
Real-world evidence (RWE) in oncology depends on patient-level electronic health record (EHR) data that is expensive and rarely open. Synthetic data generators such as Synthea offer a free alternative for developing and testing analysis pipelines. This project asks: how realistic is synthetic oncology EHR data for RWE? Specifically, if the same workflow that would be applied to a real EHR dataset is applied to a synthetic non-small cell lung cancer (NSCLC) cohort (harmonization to the OMOP Common Data Model, cohort building, survival analysis), how closely do the results agree with a population benchmark, the SEER cancer registry?
The original study design also asked how overall survival differs by stage at diagnosis and by first-line treatment. The treatment half turned out not to be answerable in this data (Section 4.5); the report therefore focuses on survival by stage and on the benchmark.
All data are synthetic or public aggregates. This is a methods demonstration, not clinical evidence, and it does not use Flatiron or any real patient-level data.
3 Data
3.1 Synthea generation and the stage patch
Patients were generated with Synthea (master-branch-latest build, 2026-08-18) with these settings from synthea/run_generate.ps1: population 60,000 living patients aged 50–90, patient seed 2026, clinician seed 2026, reference date 20261001, Massachusetts, CSV export only (claims and imaging excluded for size). The run produced 86,734 records (60,000 alive, 26,734 dead); 1,886 had a lung cancer diagnosis and were kept for the OMOP load (tools/filter_lung_cancer.py).
The stage bug. In an unmodified run every lung cancer was stage I, because the module assigns stage from a counter of undiagnosed months that almost never exceeds its first threshold. Patching that one state in lung_cancer.json was not enough: Synthea also runs veteran_lung_cancer.json for patients with the veteran attribute, with its own copy of the rule. After the first 60,000-patient run only the patched module was in place and the stage mix was still wrong; patching both modules fixed it.
The module scripts each stage’s death inside a time window after diagnosis (Stage I: 2-6 years; Stage II: 16-28 months; Stage III: 9-18 months; Stage IV: 6-10 months). This matters for the results below.
3.2 ETL-Synthea load into OMOP CDM 5.4
The Synthea CSVs were loaded into PostgreSQL 17 and mapped to OMOP CDM 5.4 with ETL-Synthea (Synthea table version 3.3.0), using the Athena vocabulary release v20260829 (vocabulary version string v5.0 29-AUG-26) (SNOMED, RxNorm, LOINC, CVX, ICDO3, Cancer Modifier and the Athena defaults). Two compatibility fixes were needed (a new NPI column in the current Synthea export, and ETL-Synthea’s loader reading ZIP codes as integers), and the OHDSI drug_era step was skipped because it did not finish in over an hour; lines of therapy use drug_exposure instead. The load was verified against the source files: persons, deaths, lung cancer conditions and the NSCLC and stage counts all match.
| OMOP table | Rows |
|---|---|
| person | 1,886 |
| visit_occurrence | 174,498 |
| condition_occurrence | 52,832 |
| drug_exposure | 496,822 |
| procedure_occurrence | 545,478 |
| measurement | 2,954,339 |
| observation | 1,059,337 |
| death | 1,818 |
| observation_period | 1,886 |
3.3 The stage III vocabulary mapping finding
In this vocabulary release the SNOMED code for “non-small cell carcinoma of lung, TNM stage 3” maps to two condition concepts: non-small cell lung cancer and small cell carcinoma of lung. ETL-Synthea writes one condition row per concept, so each of the 312 stage III patients has 2 rows, one of them small cell. As a result 576 persons carry the small cell concept: 264 true small cell patients plus the 312 stage III NSCLC patients. A cohort defined by standard concepts would silently misclassify them. Here histology and stage are therefore taken from the source codes, not the standard concepts, and a dbt test guards the rule. The standard OHDSI Data Quality Dashboard did not flag this problem (Section 5.7).
3.4 SEER benchmark
The benchmark is SEER Research Data, 17 Registries, November 2025 submission (diagnosis years 2000–2023, released April 2026), extracted with SEER*Stat: lung and bronchus, NSCLC defined by excluding ICD-O-3 histology 8041–8045 (small cell) and 8240–8249 (carcinoid and other neuroendocrine), diagnosed 2010–2015, age 50 and over, first primary only, AJCC 7th edition derived stage group, observed (all-cause) Kaplan–Meier survival in monthly intervals to December 2023. The survival cohort has 172,582 patients, 160,258 with known stage I–IV. Only aggregate results are used and published, in line with the SEER data use agreement.
4 Methods
4.1 Cohort rules (dbt)
The cohort is built with dbt models on the OMOP tables (stg_ → int_ → mart_nsclc_cohort). The index date is the first recorded lung cancer diagnosis. NSCLC versus small cell and TNM stage come from SNOMED source codes. The main cohort is staged NSCLC, age 50 or older at diagnosis (matching the SEER extraction), all diagnosis years, follow-up censored at 120 months; the sensitivity cohort is the same restricted to diagnoses in 2010–2015.
| Step | Description | Persons |
|---|---|---|
| 1 | People in the OMOP CDM (lung cancer export) | 1,886 |
| 2 | With a lung cancer diagnosis code | 1,886 |
| 3 | Excluded: small cell (source code) | 264 |
| 4 | NSCLC (source code) | 1,622 |
| 5 | Excluded: NSCLC without a single TNM stage code | 1 |
| 6 | Staged NSCLC | 1,621 |
| 7 | Excluded: age at diagnosis under 50 | 274 |
| 8 | Main analysis cohort | 1,347 |
| 9 | Sensitivity cohort (diagnosed 2010-2015) | 305 |
4.2 Endpoints
Overall survival runs from the index date to death; survivors are censored at the end of their observation period. Time to next treatment runs from the start of line 1 to the start of line 2 or death, censored at the end of observation. A prior-other-cancer flag marks patients with another cancer’s treatment or diagnosis before the NSCLC index date (flag only, nobody excluded).
4.3 Lines of therapy
Rule-based, from drug administrations of the NSCLC-directed agents (a whitelist, here cisplatin and paclitaxel) on or after diagnosis; other cancers’ regimens and hormone preparations are excluded. Line 1 starts at the first administration, with its regimen defined by the drugs started within 28 days; a gap of more than 90 days between administrations starts the next line, and a new drug alone never does. The 90-day threshold is a dbt variable, and 60 and 120 days were run as sensitivity analyses. The data support it: administrations are one day apart in 76,008 of 99,447 gaps, cycles are separated by 29–60 days (99th percentile 46 days), so a 28-day gap would start a new line in 99.6% of patients, whereas 90 days isolates 47 (3.5%).
4.4 Survival methods
Kaplan–Meier curves by stage with log(−log) confidence intervals (as SEER*Stat uses), compared with SEER at 1–5 years (SEER’s own exported estimates and confidence limits). Medians and restricted mean survival time (RMST) to 60 months, the latter for SEER computed as the area under the monthly observed-survival curve. A Cox proportional hazards model (stage + age + sex) with the proportional-hazards assumption tested by cox.zph on Schoenfeld residuals, re-fitted without patients flagged for a prior other cancer. Finally a direct-standardisation sensitivity analysis: the Synthea main cohort was reweighted to SEER’s age (50–64, 65–74, 75–84, 85+) by sex distribution within each stage, and the weighted Kaplan–Meier curve was re-estimated with a robust variance.
4.5 Why inverse probability of treatment weighting (IPTW) is not identifiable
IPTW needs variation in treatment. Among the 1,347 patients, 1,346 received the same first line (Section 5.6), so the probability of that treatment is about 1 in every stratum, propensity scores would be degenerate, and the single untreated patient would dominate any weights. A first-line treatment comparison is therefore not identifiable in this data, and no treatment-effect model is fitted.
5 Results
5.1 Cohort and demographics
| Stage | Synthea n | SEER n | % male (Synthea / SEER) | 50–64 | 65–74 | 75–84 | 85+ |
|---|---|---|---|---|---|---|---|
| I | 292 | 32,093 | 78.8 / 46.0 | 55.1 / 25.9 | 32.9 / 38.0 | 12.0 / 28.8 | 0.0 / 7.3 |
| II | 109 | 13,477 | 73.4 / 54.7 | 51.4 / 27.6 | 35.8 / 36.5 | 12.8 / 28.2 | 0.0 / 7.7 |
| III | 261 | 31,785 | 75.5 / 55.7 | 64.8 / 32.6 | 25.7 / 35.3 | 9.6 / 25.1 | 0.0 / 7.0 |
| IV | 685 | 82,903 | 76.6 / 54.8 | 58.5 / 33.7 | 29.6 / 32.8 | 11.8 / 24.6 | 0.0 / 8.9 |
| All staged (I-IV) | 1,347 | 160,258 | 76.6 / 53.2 | 58.4 / 31.4 | 30.1 / 34.6 | 11.5 / 25.8 | 0.0 / 8.1 |
The synthetic cohort is about three-quarters male (76.6% vs 53.2% in SEER) and younger (58.4% aged 50–64 vs 31.4%). Synthea has no patients aged 85 or over: its oldest patient at diagnosis is just under 85 (84.96 years), because the generator draws people aged 50–90 at the reference date, while SEER has 8.1% in that band.
5.2 Kaplan–Meier curves versus SEER
| Stage | Year | Synthea at risk | Synthea S(t), % (95% CI) | SEER S(t), % (95% CI) | Difference, pp |
|---|---|---|---|---|---|
| I | 1 | 278 | 97.2 (94.5-98.6) | 86.2 (85.8-86.6) | 11.0 |
| I | 2 | 258 | 93.7 (90.1-96.0) | 75.3 (74.9-75.8) | 18.4 |
| I | 3 | 169 | 64.3 (58.3-69.7) | 66.1 (65.6-66.6) | -1.8 |
| I | 4 | 108 | 41.4 (35.4-47.3) | 59.1 (58.5-59.6) | -17.7 |
| I | 5 | 47 | 18.2 (13.8-23.1) | 53.0 (52.5-53.6) | -34.8 |
| II | 1 | 100 | 99.0 (93.2-99.9) | 70.5 (69.7-71.3) | 28.5 |
| II | 2 | 32 | 31.7 (22.9-40.8) | 54.6 (53.8-55.4) | -22.9 |
| II | 3 | 0 | 0.0 (0.0-0.0) | 44.8 (43.9-45.6) | -44.8 |
| II | 4 | 0 | 0.0 (0.0-0.0) | 38.2 (37.3-39.0) | -38.2 |
| II | 5 | 0 | 0.0 (0.0-0.0) | 33.3 (32.5-34.1) | -33.3 |
| III | 1 | 154 | 59.7 (53.4-65.4) | 54.4 (53.8-54.9) | 5.3 |
| III | 2 | 0 | 0.0 (0.0-0.0) | 35.3 (34.8-35.8) | -35.3 |
| III | 3 | 0 | 0.0 (0.0-0.0) | 26.2 (25.7-26.7) | -26.2 |
| III | 4 | 0 | 0.0 (0.0-0.0) | 21.0 (20.5-21.4) | -21.0 |
| III | 5 | 0 | 0.0 (0.0-0.0) | 17.4 (17.0-17.8) | -17.4 |
| IV | 1 | 0 | 0.0 (0.0-0.0) | 25.0 (24.7-25.3) | -25.0 |
| IV | 2 | 0 | 0.0 (0.0-0.0) | 12.4 (12.2-12.6) | -12.4 |
| IV | 3 | 0 | 0.0 (0.0-0.0) | 7.7 (7.5-7.8) | -7.7 |
| IV | 4 | 0 | 0.0 (0.0-0.0) | 5.4 (5.2-5.5) | -5.4 |
| IV | 5 | 0 | 0.0 (0.0-0.0) | 4.1 (3.9-4.2) | -4.1 |
5.3 Medians and restricted mean survival time
| Stage | Median, Synthea (95% CI) | Median, SEER | RMST 60 mo, Synthea (95% CI) | RMST 60 mo, SEER | RMST difference |
|---|---|---|---|---|---|
| I | 43.8 (41.3–47.0) | 67 | 42.9 (41.3-44.6) | 43.45 | -0.54 |
| II | 21.4 (20.6–22.9) | 29 | 21.6 (20.8-22.3) | 32.57 | -11.00 |
| III | 13.2 (12.5–14.0) | 14 | 13.3 (13.0-13.6) | 22.75 | -9.45 |
| IV | 8.0 (7.9–8.2) | 5 | 8.0 (7.9-8.1) | 10.27 | -2.28 |
The summaries agree for some stages (the stage III median, the stage I RMST) and not for others: the stage I median is 43.8 vs 67 months, stage II 21.4 vs 29, and the stage II and III RMSTs are about 11 and 9 months shorter than SEER’s. A single number can look right while the curve is wrong.
5.4 Cox model and the proportional-hazards test
| Term | Hazard ratio | 95% CI | p |
|---|---|---|---|
| Stage II (vs I) | 22.2 | 14.7 – 33.4 | <1e-16 |
| Stage III (vs I) | 290 | 170 – 494 | <1e-16 |
| Stage IV (vs I) | 9,490 | 5,060 – 17,800 | <1e-16 |
| Age (per year) | 1.005 | 0.998 – 1.011 | 0.17 |
| Male (vs female) | 1.01 | 0.886 – 1.16 | 0.84 |
| Age (per 10 years) | 1.05 | 0.98 – 1.12 | 0.17 |
The stage hazard ratios (22, 290 and 9,488 against stage I) reflect near-complete separation of survival times by stage, not clinical effect sizes. Each stage’s death falls inside a scripted window, and neighbouring windows overlap only at their edges, so the model is close to separation and its estimates and intervals are numerically extreme. The proportional-hazards assumption fails for stage (global test p < 1e-16; by coefficient, stage II p = 0.0025, stage III p = 9.2e-05), consistent with the Kaplan–Meier comparison. Age (HR 1.05 per 10 years) and sex (HR 1.01) have essentially no effect in Synthea, whereas in real NSCLC observed survival worsens with age and is generally lower in men: the simulator’s survival after diagnosis depends on stage alone. Excluding the 241 patients flagged for a prior other cancer does not change this pattern.
5.5 Reweighting to SEER’s age and sex distribution
| Weighting | Stage | Unweighted, % (95% CI) | Reweighted, % (95% CI) | Change, pp | SEER, % |
|---|---|---|---|---|---|
| 4 age bands | I | 18.2 (13.8-23.1) | 19.0 (11.6-27.7) | 0.7 | 53.0 |
| 4 age bands | II | 0.0 (0.0-0.0) | 0.0 (0.0-0.0) | 0.0 | 33.3 |
| 4 age bands | III | 0.0 (0.0-0.0) | 0.0 (0.0-0.0) | 0.0 | 17.4 |
| 4 age bands | IV | 0.0 (0.0-0.0) | 0.0 (0.0-0.0) | 0.0 | 4.1 |
| 3 age bands (75+ merged) | I | 18.2 (13.8-23.1) | 17.8 (10.7-26.4) | -0.4 | 53.0 |
| 3 age bands (75+ merged) | II | 0.0 (0.0-0.0) | 0.0 (0.0-0.0) | 0.0 | 33.3 |
| 3 age bands (75+ merged) | III | 0.0 (0.0-0.0) | 0.0 (0.0-0.0) | 0.0 | 17.4 |
| 3 age bands (75+ merged) | IV | 0.0 (0.0-0.0) | 0.0 (0.0-0.0) | 0.0 | 4.1 |
Reweighting barely moves the five-year estimates: stage I goes from 18.2% to 17.8–19.0% against 53.0% in SEER, and stages II–IV stay at 0%, because every such patient has died long before five years whatever their age or sex. Demographic differences therefore do not explain the gap with SEER; some intermediate years move by several percentage points, with wide intervals, in strata with few patients at risk (full table in results/km_synthea_vs_seer_reweighted.csv).
5.6 Treatment patterns and lines of therapy
| Stage | chemotherapy + radiation | chemotherapy alone | none recorded |
|---|---|---|---|
| I | 292 | 0 | 0 |
| II | 109 | 0 | 0 |
| III | 261 | 0 | 0 |
| IV | 684 | 0 | 1 |
1,346 of 1,347 patients received cisplatin and paclitaxel together with the combined chemotherapy and radiation procedure, a median 2 days after diagnosis (IQR 2–3); one patient has no recorded treatment. First-line type is not even associated with stage (Cramér’s V 0.027, p = 0.81): it is constant. The number of administration dates per patient rises with survival (median 181, 94, 56 and 32 in stages I to IV), because treatment simply continues until death.
| Gap, days | Treated patients | With a line 2 | With 3 or more lines | Total lines |
|---|---|---|---|---|
| 60 | 1,346 | 168 | 32 | 1,559 |
| 90 | 1,346 | 47 | 1 | 1,394 |
| 120 | 1,346 | 20 | 0 | 1,366 |
At the default gap 47 patients (3.5%) have a second line, and every line 1 and line 2 is cisplatin plus paclitaxel. Time to next treatment is therefore essentially time to death: of 1,293 events only 47 are line 2 starts (median 9.4 months). 241 patients (17.9%) had another cancer’s treatment or diagnosis before the NSCLC diagnosis (199 by drug, 30 by procedure, 73 by diagnosis; overlapping).
5.7 Data quality (OHDSI Data Quality Dashboard)
| Category | Checks | Passed | Failed | Not applicable | Errored |
|---|---|---|---|---|---|
| Completeness | 501 | 354 | 8 | 128 | 11 |
| Conformance | 1060 | 815 | 4 | 223 | 18 |
| Plausibility | 813 | 321 | 6 | 482 | 4 |
Of 2,374 checks, 1,490 passed and 18 failed (833 were not applicable because the table is empty and 33 errored because the optional COHORT tables do not exist). Two of the failed checks are on those absent tables, so 16 are threshold failures:
| Category | Check | Table.field | % violated | Cause |
|---|---|---|---|---|
| Completeness | measurePersonCompleteness | DRUG_ERA | 100.0 | ETL (deliberate: step skipped) |
| Conformance | fkClass | DRUG_STRENGTH.INGREDIENT_CONCEPT_ID | 0.1 | vocabulary release |
| Conformance | isStandardValidConcept | OBSERVATION.OBSERVATION_TYPE_CONCEPT_ID | 99.9 | ETL |
| Completeness | standardConceptRecordCompleteness | CONDITION_OCCURRENCE.CONDITION_STATUS_CONCEPT_ID | 100.0 | ETL |
| Completeness | standardConceptRecordCompleteness | MEASUREMENT.UNIT_CONCEPT_ID | 37.0 | Synthea |
| Completeness | standardConceptRecordCompleteness | OBSERVATION.UNIT_CONCEPT_ID | 100.0 | ETL |
| Completeness | standardConceptRecordCompleteness | VISIT_DETAIL.ADMITTED_FROM_CONCEPT_ID | 100.0 | Synthea |
| Completeness | standardConceptRecordCompleteness | VISIT_DETAIL.DISCHARGED_TO_CONCEPT_ID | 100.0 | Synthea |
| Completeness | standardConceptRecordCompleteness | VISIT_OCCURRENCE.ADMITTED_FROM_CONCEPT_ID | 100.0 | Synthea |
| Completeness | standardConceptRecordCompleteness | VISIT_OCCURRENCE.DISCHARGED_TO_CONCEPT_ID | 100.0 | Synthea |
| Plausibility | plausibleValueLow | DRUG_EXPOSURE.DAYS_SUPPLY | 75.8 | ETL / Synthea |
| Plausibility | plausibleValueLow | DRUG_EXPOSURE.QUANTITY | 100.0 | ETL |
| Plausibility | plausibleValueLow | OBSERVATION_PERIOD.OBSERVATION_PERIOD_START_DATE | 13.4 | Synthea |
| Plausibility | plausibleValueHigh | DRUG_EXPOSURE.DAYS_SUPPLY | 2.6 | ETL / Synthea |
| Plausibility | plausibleBeforeDeath | PAYER_PLAN_PERIOD.PAYER_PLAN_PERIOD_END_DATE | 2.0 | Synthea |
| Plausibility | plausibleUnitConceptIds | MEASUREMENT.MEASUREMENT_CONCEPT_ID | 91.8 | Synthea |
Two failures touch the cohort’s source tables (CONDITION_OCCURRENCE.CONDITION_STATUS_CONCEPT_ID is empty and OBSERVATION_PERIOD starts before 1950 for older patients), and neither affects a field the cohort uses; every check on person, death and the cohort’s condition and observation-period fields passes. The dashboard did not detect the stage III small-cell mapping: the extra rows are well-formed and plausible, which is why that problem was found by cohort-level tests and by comparing the load with the source files.
6 Discussion
What Synthea gets right. It orders the stages correctly (median survival falls from stage I to IV), its stage mix can be calibrated to SEER, and it provides the structure a real EHR pipeline needs: encounters, conditions, drugs, procedures and deaths that map cleanly into OMOP, with realistic messiness such as other cancers, hormone preparations and missing fields. As a test bed for the pipeline (ETL, cohort definitions, tests, survival code) it did its job, including exposing a vocabulary mapping problem.
What it gets wrong. The shape of survival is wrong. Deaths are scripted in stage-specific windows, so survival is flat and then collapses, nobody with stage II–IV disease is followed beyond 28 months, and the curves cannot be reconciled with SEER’s long tails. There is no age or sex effect. Treatment is constant, so the treatment question cannot be studied at all. Demographics differ (three-quarters male, no patients aged 85 or over). Calibrating the stage mix to SEER makes the stage distribution agree by construction, but not the outcomes.
Lesson: validate full distributions, not summary statistics. A stage III median of 13.2 vs 14 months and a stage I RMST within about a month of SEER’s would have passed this data as realistic. Only the full curves, the five-year survival, the proportional-hazards test and the treatment table show that it is not. A validation of synthetic data against a benchmark should compare whole curves, subgroup by subgroup, and check what the data generator cannot vary.
7 Limitations
- The data are synthetic: effect estimates are not clinical evidence, and the Synthea lung cancer module scripts survival by design.
- One simulated run (one seed) was analysed; results are not averaged over simulations.
- SEER survival is observed (all-cause) survival, includes first primary cancers only, and uses AJCC 7th edition stage; the Synthea stages are TNM labels mapped to the same four groups. Calendar-period effects (such as COVID-19-era mortality) and age-related background mortality affect SEER’s observed survival; in Synthea, age has no detectable effect on survival (Section 5.4).
- SEER intervals at 1–5 years are SEER*Stat’s own log(−log) limits; the survival curves of the two sources are compared at yearly points, not by a formal test.
- The stage mix was set from SEER, so agreement on stage distribution is by construction; the benchmark is informative about survival, not stage mix.
- The sensitivity cohort (diagnosed 2010–2015) is small (305 patients).
- Reweighting could not represent SEER’s 85+ patients (no Synthea patients), and its effective sample size is much smaller than the cohort.
- The Cox hazard ratios are numerically extreme (near-separation) and the proportional-hazards assumption fails for stage.
- The vocabulary findings (stage III mapping, DQD observations) refer to one Athena release and one ETL version.
- Time to next treatment is mostly time to death because second lines are rare.
8 Reproducibility
The repository is at https://github.com/erickyegon/oncology-rwe-nsclc and contains all code; the generated data are not committed. To reproduce (from the repository root; prerequisites and setup are in the README):
- Create the database and role:
etl/setup_postgres.ps1; generate the data:synthea/run_generate.ps1; keep lung cancer patients:tools/filter_lung_cancer.py. - Load the vocabularies and the Synthea tables and map to OMOP:
etl/01_load_vocab.R,etl/01b_load_vocab_copy.sh,etl/01c_vocab_indexes.sql,etl/02a_load_native_copy.sh, thenRscript etl/02_etl_synthea.R; check the load withtools/verify_omop.py. - Build and test the cohort:
dbt build --project-dir dbt --profiles-dir dbt(profile template indbt/profiles.example.yml). - Analyses:
analysis/km_synthea_vs_seer.R,analysis/cox_os.R,analysis/rmst.R,analysis/seer_age_sex_reweighting.Randanalysis/report_inputs.R. The SEER-based scripts need the local SEER*Stat exports (not in the repository); the committed aggregates inseer/andresults/are enough to render this report. - Render this report:
quarto render report/report.qmd(every number is read fromresults/,seer/and the repository files).
Software: R 4.6.1, PostgreSQL 17, dbt-postgres 1.12, Quarto 1.10, OHDSI ETL-Synthea 2.1, DataQualityDashboard 2.8.9. Library versions are recorded in results/sessionInfo.txt.
9 Use of AI tools
I used AI assistants (Claude) as tools in this project: to write and debug code, run checks, and draft documentation. The research question, study design, cohort and NSCLC definitions, SEER extraction, methodological choices and interpretation of the results are my own. I reviewed every output, verified results against their sources (including the SEER*Stat exports and the OMOP source data), and I am responsible for all content in this repository. Commits made with AI assistance carry Co-Authored-By trailers.
Code and results: https://github.com/erickyegon/oncology-rwe-nsclc. Medium article.