How realistic is synthetic oncology EHR data?

NSCLC survival from Synthea, harmonized to OMOP and benchmarked against SEER

Author

Erick Kiprotich Yegon (linkedin.com/in/erickyegon)

Published

October 2, 2026

1 Summary

This project built an end-to-end workflow (Synthea → OMOP CDM 5.4 → dbt → R) on a synthetic NSCLC cohort of 1,347 patients and benchmarked its survival against 172,582 real U.S. patients from SEER (160,258 with known stage). Unmodified Synthea staged every lung cancer as stage I, and even after patching that module a second, hidden path (the veteran module) still produced 51.5% stage I, so both had to be fixed before any comparison was possible; the stage mix now matches SEER (chi-square p = 0.50 for age 50 and over). Synthea ranks the stages correctly, and some single summaries look close to SEER, such as the stage III median (13.2 vs 14 months) and the stage I restricted mean survival time (42.9 vs 43.5 months), yet the curves do not match: nobody with stage II–IV disease is followed beyond 28 months, whereas SEER five-year survival is 33%, 17% and 4% for stages II, III and IV. Age and sex have no effect on survival in Synthea (hazard ratio 1.05 per 10 years and 1.01 for men), reweighting to SEER’s age and sex distribution leaves the curves essentially unchanged, and 1,346 of 1,347 patients received the same treatment, so treatment effects cannot be identified. Synthetic data can test a pipeline but not stand in for real-world outcomes, and the lesson is to validate full distributions, not summary statistics.

2 Background and question

Real-world evidence (RWE) in oncology depends on patient-level electronic health record (EHR) data that is expensive and rarely open. Synthetic data generators such as Synthea offer a free alternative for developing and testing analysis pipelines. This project asks: how realistic is synthetic oncology EHR data for RWE? Specifically, if the same workflow that would be applied to a real EHR dataset is applied to a synthetic non-small cell lung cancer (NSCLC) cohort (harmonization to the OMOP Common Data Model, cohort building, survival analysis), how closely do the results agree with a population benchmark, the SEER cancer registry?

The original study design also asked how overall survival differs by stage at diagnosis and by first-line treatment. The treatment half turned out not to be answerable in this data (Section 4.5); the report therefore focuses on survival by stage and on the benchmark.

All data are synthetic or public aggregates. This is a methods demonstration, not clinical evidence, and it does not use Flatiron or any real patient-level data.

3 Data

3.1 Synthea generation and the stage patch

Patients were generated with Synthea (master-branch-latest build, 2026-08-18) with these settings from synthea/run_generate.ps1: population 60,000 living patients aged 50–90, patient seed 2026, clinician seed 2026, reference date 20261001, Massachusetts, CSV export only (claims and imaging excluded for size). The run produced 86,734 records (60,000 alive, 26,734 dead); 1,886 had a lung cancer diagnosis and were kept for the OMOP load (tools/filter_lung_cancer.py).

The stage bug. In an unmodified run every lung cancer was stage I, because the module assigns stage from a counter of undiagnosed months that almost never exceeds its first threshold. Patching that one state in lung_cancer.json was not enough: Synthea also runs veteran_lung_cancer.json for patients with the veteran attribute, with its own copy of the rule. After the first 60,000-patient run only the patched module was in place and the stage mix was still wrong; patching both modules fixed it.

Stage mix of staged NSCLC patients (all ages): the SEER input, the shares used in the patched modules, and the Synthea runs before and after patching the second module. SEER n = 160,258 with known stage.
Stage SEER stage mix (known stage), % Share written into both modules, % First run (veteran module unpatched), % Analysis run (both patched), %
I 20.0 20.02 51.5 21.8
II 8.4 8.41 5.8 7.4
III 19.8 19.82 11.6 19.2
IV 51.7 51.75 31.1 51.5

The module scripts each stage’s death inside a time window after diagnosis (Stage I: 2-6 years; Stage II: 16-28 months; Stage III: 9-18 months; Stage IV: 6-10 months). This matters for the results below.

3.2 ETL-Synthea load into OMOP CDM 5.4

The Synthea CSVs were loaded into PostgreSQL 17 and mapped to OMOP CDM 5.4 with ETL-Synthea (Synthea table version 3.3.0), using the Athena vocabulary release v20260829 (vocabulary version string v5.0 29-AUG-26) (SNOMED, RxNorm, LOINC, CVX, ICDO3, Cancer Modifier and the Athena defaults). Two compatibility fixes were needed (a new NPI column in the current Synthea export, and ETL-Synthea’s loader reading ZIP codes as integers), and the OHDSI drug_era step was skipped because it did not finish in over an hour; lines of therapy use drug_exposure instead. The load was verified against the source files: persons, deaths, lung cancer conditions and the NSCLC and stage counts all match.

Rows loaded into the OMOP CDM (lung cancer patients only).
OMOP table Rows
person 1,886
visit_occurrence 174,498
condition_occurrence 52,832
drug_exposure 496,822
procedure_occurrence 545,478
measurement 2,954,339
observation 1,059,337
death 1,818
observation_period 1,886

3.3 The stage III vocabulary mapping finding

In this vocabulary release the SNOMED code for “non-small cell carcinoma of lung, TNM stage 3” maps to two condition concepts: non-small cell lung cancer and small cell carcinoma of lung. ETL-Synthea writes one condition row per concept, so each of the 312 stage III patients has 2 rows, one of them small cell. As a result 576 persons carry the small cell concept: 264 true small cell patients plus the 312 stage III NSCLC patients. A cohort defined by standard concepts would silently misclassify them. Here histology and stage are therefore taken from the source codes, not the standard concepts, and a dbt test guards the rule. The standard OHDSI Data Quality Dashboard did not flag this problem (Section 5.7).

3.4 SEER benchmark

The benchmark is SEER Research Data, 17 Registries, November 2025 submission (diagnosis years 2000–2023, released April 2026), extracted with SEER*Stat: lung and bronchus, NSCLC defined by excluding ICD-O-3 histology 8041–8045 (small cell) and 8240–8249 (carcinoid and other neuroendocrine), diagnosed 2010–2015, age 50 and over, first primary only, AJCC 7th edition derived stage group, observed (all-cause) Kaplan–Meier survival in monthly intervals to December 2023. The survival cohort has 172,582 patients, 160,258 with known stage I–IV. Only aggregate results are used and published, in line with the SEER data use agreement.

4 Methods

4.1 Cohort rules (dbt)

The cohort is built with dbt models on the OMOP tables (stg_ → int_ → mart_nsclc_cohort). The index date is the first recorded lung cancer diagnosis. NSCLC versus small cell and TNM stage come from SNOMED source codes. The main cohort is staged NSCLC, age 50 or older at diagnosis (matching the SEER extraction), all diagnosis years, follow-up censored at 120 months; the sensitivity cohort is the same restricted to diagnoses in 2010–2015.

Cohort flow, from the dbt model mart_cohort_attrition.
Step Description Persons
1 People in the OMOP CDM (lung cancer export) 1,886
2 With a lung cancer diagnosis code 1,886
3 Excluded: small cell (source code) 264
4 NSCLC (source code) 1,622
5 Excluded: NSCLC without a single TNM stage code 1
6 Staged NSCLC 1,621
7 Excluded: age at diagnosis under 50 274
8 Main analysis cohort 1,347
9 Sensitivity cohort (diagnosed 2010-2015) 305

4.2 Endpoints

Overall survival runs from the index date to death; survivors are censored at the end of their observation period. Time to next treatment runs from the start of line 1 to the start of line 2 or death, censored at the end of observation. A prior-other-cancer flag marks patients with another cancer’s treatment or diagnosis before the NSCLC index date (flag only, nobody excluded).

4.3 Lines of therapy

Rule-based, from drug administrations of the NSCLC-directed agents (a whitelist, here cisplatin and paclitaxel) on or after diagnosis; other cancers’ regimens and hormone preparations are excluded. Line 1 starts at the first administration, with its regimen defined by the drugs started within 28 days; a gap of more than 90 days between administrations starts the next line, and a new drug alone never does. The 90-day threshold is a dbt variable, and 60 and 120 days were run as sensitivity analyses. The data support it: administrations are one day apart in 76,008 of 99,447 gaps, cycles are separated by 29–60 days (99th percentile 46 days), so a 28-day gap would start a new line in 99.6% of patients, whereas 90 days isolates 47 (3.5%).

4.4 Survival methods

Kaplan–Meier curves by stage with log(−log) confidence intervals (as SEER*Stat uses), compared with SEER at 1–5 years (SEER’s own exported estimates and confidence limits). Medians and restricted mean survival time (RMST) to 60 months, the latter for SEER computed as the area under the monthly observed-survival curve. A Cox proportional hazards model (stage + age + sex) with the proportional-hazards assumption tested by cox.zph on Schoenfeld residuals, re-fitted without patients flagged for a prior other cancer. Finally a direct-standardisation sensitivity analysis: the Synthea main cohort was reweighted to SEER’s age (50–64, 65–74, 75–84, 85+) by sex distribution within each stage, and the weighted Kaplan–Meier curve was re-estimated with a robust variance.

4.5 Why inverse probability of treatment weighting (IPTW) is not identifiable

IPTW needs variation in treatment. Among the 1,347 patients, 1,346 received the same first line (Section 5.6), so the probability of that treatment is about 1 in every stratum, propensity scores would be degenerate, and the single untreated patient would dominate any weights. A first-line treatment comparison is therefore not identifiable in this data, and no treatment-effect model is fitted.

5 Results

5.1 Cohort and demographics

Synthea main cohort vs SEER: percent male and percent in each age band at diagnosis (Synthea / SEER).
Stage Synthea n SEER n % male (Synthea / SEER) 50–64 65–74 75–84 85+
I 292 32,093 78.8 / 46.0 55.1 / 25.9 32.9 / 38.0 12.0 / 28.8 0.0 / 7.3
II 109 13,477 73.4 / 54.7 51.4 / 27.6 35.8 / 36.5 12.8 / 28.2 0.0 / 7.7
III 261 31,785 75.5 / 55.7 64.8 / 32.6 25.7 / 35.3 9.6 / 25.1 0.0 / 7.0
IV 685 82,903 76.6 / 54.8 58.5 / 33.7 29.6 / 32.8 11.8 / 24.6 0.0 / 8.9
All staged (I-IV) 1,347 160,258 76.6 / 53.2 58.4 / 31.4 30.1 / 34.6 11.5 / 25.8 0.0 / 8.1

The synthetic cohort is about three-quarters male (76.6% vs 53.2% in SEER) and younger (58.4% aged 50–64 vs 31.4%). Synthea has no patients aged 85 or over: its oldest patient at diagnosis is just under 85 (84.96 years), because the generator draws people aged 50–90 at the reference date, while SEER has 8.1% in that band.

5.2 Kaplan–Meier curves versus SEER

Kaplan–Meier survival by stage in the Synthea main cohort (steps, 95% log-log CI, numbers at risk) against SEER observed survival at years 1–5.
Survival at 1–5 years, main cohort. S(t) is 0 once every patient in a stage has died (no one is at risk).
Stage Year Synthea at risk Synthea S(t), % (95% CI) SEER S(t), % (95% CI) Difference, pp
I 1 278 97.2 (94.5-98.6) 86.2 (85.8-86.6) 11.0
I 2 258 93.7 (90.1-96.0) 75.3 (74.9-75.8) 18.4
I 3 169 64.3 (58.3-69.7) 66.1 (65.6-66.6) -1.8
I 4 108 41.4 (35.4-47.3) 59.1 (58.5-59.6) -17.7
I 5 47 18.2 (13.8-23.1) 53.0 (52.5-53.6) -34.8
II 1 100 99.0 (93.2-99.9) 70.5 (69.7-71.3) 28.5
II 2 32 31.7 (22.9-40.8) 54.6 (53.8-55.4) -22.9
II 3 0 0.0 (0.0-0.0) 44.8 (43.9-45.6) -44.8
II 4 0 0.0 (0.0-0.0) 38.2 (37.3-39.0) -38.2
II 5 0 0.0 (0.0-0.0) 33.3 (32.5-34.1) -33.3
III 1 154 59.7 (53.4-65.4) 54.4 (53.8-54.9) 5.3
III 2 0 0.0 (0.0-0.0) 35.3 (34.8-35.8) -35.3
III 3 0 0.0 (0.0-0.0) 26.2 (25.7-26.7) -26.2
III 4 0 0.0 (0.0-0.0) 21.0 (20.5-21.4) -21.0
III 5 0 0.0 (0.0-0.0) 17.4 (17.0-17.8) -17.4
IV 1 0 0.0 (0.0-0.0) 25.0 (24.7-25.3) -25.0
IV 2 0 0.0 (0.0-0.0) 12.4 (12.2-12.6) -12.4
IV 3 0 0.0 (0.0-0.0) 7.7 (7.5-7.8) -7.7
IV 4 0 0.0 (0.0-0.0) 5.4 (5.2-5.5) -5.4
IV 5 0 0.0 (0.0-0.0) 4.1 (3.9-4.2) -4.1

5.3 Medians and restricted mean survival time

Median survival and restricted mean survival time to 60 months (months), main cohort.
Stage Median, Synthea (95% CI) Median, SEER RMST 60 mo, Synthea (95% CI) RMST 60 mo, SEER RMST difference
I 43.8 (41.3–47.0) 67 42.9 (41.3-44.6) 43.45 -0.54
II 21.4 (20.6–22.9) 29 21.6 (20.8-22.3) 32.57 -11.00
III 13.2 (12.5–14.0) 14 13.3 (13.0-13.6) 22.75 -9.45
IV 8.0 (7.9–8.2) 5 8.0 (7.9-8.1) 10.27 -2.28

The summaries agree for some stages (the stage III median, the stage I RMST) and not for others: the stage I median is 43.8 vs 67 months, stage II 21.4 vs 29, and the stage II and III RMSTs are about 11 and 9 months shorter than SEER’s. A single number can look right while the curve is wrong.

5.4 Cox model and the proportional-hazards test

Cox model for overall survival, main cohort (reference: stage I, female).
Term Hazard ratio 95% CI p
Stage II (vs I) 22.2 14.7 – 33.4 <1e-16
Stage III (vs I) 290 170 – 494 <1e-16
Stage IV (vs I) 9,490 5,060 – 17,800 <1e-16
Age (per year) 1.005 0.998 – 1.011 0.17
Male (vs female) 1.01 0.886 – 1.16 0.84
Age (per 10 years) 1.05 0.98 – 1.12 0.17

The stage hazard ratios (22, 290 and 9,488 against stage I) reflect near-complete separation of survival times by stage, not clinical effect sizes. Each stage’s death falls inside a scripted window, and neighbouring windows overlap only at their edges, so the model is close to separation and its estimates and intervals are numerically extreme. The proportional-hazards assumption fails for stage (global test p < 1e-16; by coefficient, stage II p = 0.0025, stage III p = 9.2e-05), consistent with the Kaplan–Meier comparison. Age (HR 1.05 per 10 years) and sex (HR 1.01) have essentially no effect in Synthea, whereas in real NSCLC observed survival worsens with age and is generally lower in men: the simulator’s survival after diagnosis depends on stage alone. Excluding the 241 patients flagged for a prior other cancer does not change this pattern.

Schoenfeld residuals by coefficient, main cohort.

5.5 Reweighting to SEER’s age and sex distribution

Five-year survival before and after reweighting the Synthea main cohort to SEER’s age by sex distribution within each stage. The 4-band weighting drops SEER’s 85+ cells, which have no Synthea patients; the 3-band weighting merges ages 75 and over.
Weighting Stage Unweighted, % (95% CI) Reweighted, % (95% CI) Change, pp SEER, %
4 age bands I 18.2 (13.8-23.1) 19.0 (11.6-27.7) 0.7 53.0
4 age bands II 0.0 (0.0-0.0) 0.0 (0.0-0.0) 0.0 33.3
4 age bands III 0.0 (0.0-0.0) 0.0 (0.0-0.0) 0.0 17.4
4 age bands IV 0.0 (0.0-0.0) 0.0 (0.0-0.0) 0.0 4.1
3 age bands (75+ merged) I 18.2 (13.8-23.1) 17.8 (10.7-26.4) -0.4 53.0
3 age bands (75+ merged) II 0.0 (0.0-0.0) 0.0 (0.0-0.0) 0.0 33.3
3 age bands (75+ merged) III 0.0 (0.0-0.0) 0.0 (0.0-0.0) 0.0 17.4
3 age bands (75+ merged) IV 0.0 (0.0-0.0) 0.0 (0.0-0.0) 0.0 4.1

Reweighting barely moves the five-year estimates: stage I goes from 18.2% to 17.8–19.0% against 53.0% in SEER, and stages II–IV stay at 0%, because every such patient has died long before five years whatever their age or sex. Demographic differences therefore do not explain the gap with SEER; some intermediate years move by several percentage points, with wide intervals, in strata with few patients at risk (full table in results/km_synthea_vs_seer_reweighted.csv).

5.6 Treatment patterns and lines of therapy

First-line treatment type by stage, main cohort (chemoradiation if the combined procedure falls within 28 days of the first systemic dose).
Stage chemotherapy + radiation chemotherapy alone none recorded
I 292 0 0
II 109 0 0
III 261 0 0
IV 684 0 1

1,346 of 1,347 patients received cisplatin and paclitaxel together with the combined chemotherapy and radiation procedure, a median 2 days after diagnosis (IQR 2–3); one patient has no recorded treatment. First-line type is not even associated with stage (Cramér’s V 0.027, p = 0.81): it is constant. The number of administration dates per patient rises with survival (median 181, 94, 56 and 32 in stages I to IV), because treatment simply continues until death.

Lines of therapy by gap length (the default is 90 days).
Gap, days Treated patients With a line 2 With 3 or more lines Total lines
60 1,346 168 32 1,559
90 1,346 47 1 1,394
120 1,346 20 0 1,366

At the default gap 47 patients (3.5%) have a second line, and every line 1 and line 2 is cisplatin plus paclitaxel. Time to next treatment is therefore essentially time to death: of 1,293 events only 47 are line 2 starts (median 9.4 months). 241 patients (17.9%) had another cancer’s treatment or diagnosis before the NSCLC diagnosis (199 by drug, 30 by procedure, 73 by diagnosis; overlapping).

5.7 Data quality (OHDSI Data Quality Dashboard)

OHDSI Data Quality Dashboard, CDM 5.4: checks by category.
Category Checks Passed Failed Not applicable Errored
Completeness 501 354 8 128 11
Conformance 1060 815 4 223 18
Plausibility 813 321 6 482 4

Of 2,374 checks, 1,490 passed and 18 failed (833 were not applicable because the table is empty and 33 errored because the optional COHORT tables do not exist). Two of the failed checks are on those absent tables, so 16 are threshold failures:

Threshold failures and their causes (Synthea artefact, ETL, or vocabulary release).
Category Check Table.field % violated Cause
Completeness measurePersonCompleteness DRUG_ERA 100.0 ETL (deliberate: step skipped)
Conformance fkClass DRUG_STRENGTH.INGREDIENT_CONCEPT_ID 0.1 vocabulary release
Conformance isStandardValidConcept OBSERVATION.OBSERVATION_TYPE_CONCEPT_ID 99.9 ETL
Completeness standardConceptRecordCompleteness CONDITION_OCCURRENCE.CONDITION_STATUS_CONCEPT_ID 100.0 ETL
Completeness standardConceptRecordCompleteness MEASUREMENT.UNIT_CONCEPT_ID 37.0 Synthea
Completeness standardConceptRecordCompleteness OBSERVATION.UNIT_CONCEPT_ID 100.0 ETL
Completeness standardConceptRecordCompleteness VISIT_DETAIL.ADMITTED_FROM_CONCEPT_ID 100.0 Synthea
Completeness standardConceptRecordCompleteness VISIT_DETAIL.DISCHARGED_TO_CONCEPT_ID 100.0 Synthea
Completeness standardConceptRecordCompleteness VISIT_OCCURRENCE.ADMITTED_FROM_CONCEPT_ID 100.0 Synthea
Completeness standardConceptRecordCompleteness VISIT_OCCURRENCE.DISCHARGED_TO_CONCEPT_ID 100.0 Synthea
Plausibility plausibleValueLow DRUG_EXPOSURE.DAYS_SUPPLY 75.8 ETL / Synthea
Plausibility plausibleValueLow DRUG_EXPOSURE.QUANTITY 100.0 ETL
Plausibility plausibleValueLow OBSERVATION_PERIOD.OBSERVATION_PERIOD_START_DATE 13.4 Synthea
Plausibility plausibleValueHigh DRUG_EXPOSURE.DAYS_SUPPLY 2.6 ETL / Synthea
Plausibility plausibleBeforeDeath PAYER_PLAN_PERIOD.PAYER_PLAN_PERIOD_END_DATE 2.0 Synthea
Plausibility plausibleUnitConceptIds MEASUREMENT.MEASUREMENT_CONCEPT_ID 91.8 Synthea

Two failures touch the cohort’s source tables (CONDITION_OCCURRENCE.CONDITION_STATUS_CONCEPT_ID is empty and OBSERVATION_PERIOD starts before 1950 for older patients), and neither affects a field the cohort uses; every check on person, death and the cohort’s condition and observation-period fields passes. The dashboard did not detect the stage III small-cell mapping: the extra rows are well-formed and plausible, which is why that problem was found by cohort-level tests and by comparing the load with the source files.

6 Discussion

What Synthea gets right. It orders the stages correctly (median survival falls from stage I to IV), its stage mix can be calibrated to SEER, and it provides the structure a real EHR pipeline needs: encounters, conditions, drugs, procedures and deaths that map cleanly into OMOP, with realistic messiness such as other cancers, hormone preparations and missing fields. As a test bed for the pipeline (ETL, cohort definitions, tests, survival code) it did its job, including exposing a vocabulary mapping problem.

What it gets wrong. The shape of survival is wrong. Deaths are scripted in stage-specific windows, so survival is flat and then collapses, nobody with stage II–IV disease is followed beyond 28 months, and the curves cannot be reconciled with SEER’s long tails. There is no age or sex effect. Treatment is constant, so the treatment question cannot be studied at all. Demographics differ (three-quarters male, no patients aged 85 or over). Calibrating the stage mix to SEER makes the stage distribution agree by construction, but not the outcomes.

Lesson: validate full distributions, not summary statistics. A stage III median of 13.2 vs 14 months and a stage I RMST within about a month of SEER’s would have passed this data as realistic. Only the full curves, the five-year survival, the proportional-hazards test and the treatment table show that it is not. A validation of synthetic data against a benchmark should compare whole curves, subgroup by subgroup, and check what the data generator cannot vary.

7 Limitations

  • The data are synthetic: effect estimates are not clinical evidence, and the Synthea lung cancer module scripts survival by design.
  • One simulated run (one seed) was analysed; results are not averaged over simulations.
  • SEER survival is observed (all-cause) survival, includes first primary cancers only, and uses AJCC 7th edition stage; the Synthea stages are TNM labels mapped to the same four groups. Calendar-period effects (such as COVID-19-era mortality) and age-related background mortality affect SEER’s observed survival; in Synthea, age has no detectable effect on survival (Section 5.4).
  • SEER intervals at 1–5 years are SEER*Stat’s own log(−log) limits; the survival curves of the two sources are compared at yearly points, not by a formal test.
  • The stage mix was set from SEER, so agreement on stage distribution is by construction; the benchmark is informative about survival, not stage mix.
  • The sensitivity cohort (diagnosed 2010–2015) is small (305 patients).
  • Reweighting could not represent SEER’s 85+ patients (no Synthea patients), and its effective sample size is much smaller than the cohort.
  • The Cox hazard ratios are numerically extreme (near-separation) and the proportional-hazards assumption fails for stage.
  • The vocabulary findings (stage III mapping, DQD observations) refer to one Athena release and one ETL version.
  • Time to next treatment is mostly time to death because second lines are rare.

8 Reproducibility

The repository is at https://github.com/erickyegon/oncology-rwe-nsclc and contains all code; the generated data are not committed. To reproduce (from the repository root; prerequisites and setup are in the README):

  1. Create the database and role: etl/setup_postgres.ps1; generate the data: synthea/run_generate.ps1; keep lung cancer patients: tools/filter_lung_cancer.py.
  2. Load the vocabularies and the Synthea tables and map to OMOP: etl/01_load_vocab.R, etl/01b_load_vocab_copy.sh, etl/01c_vocab_indexes.sql, etl/02a_load_native_copy.sh, then Rscript etl/02_etl_synthea.R; check the load with tools/verify_omop.py.
  3. Build and test the cohort: dbt build --project-dir dbt --profiles-dir dbt (profile template in dbt/profiles.example.yml).
  4. Analyses: analysis/km_synthea_vs_seer.R, analysis/cox_os.R, analysis/rmst.R, analysis/seer_age_sex_reweighting.R and analysis/report_inputs.R. The SEER-based scripts need the local SEER*Stat exports (not in the repository); the committed aggregates in seer/ and results/ are enough to render this report.
  5. Render this report: quarto render report/report.qmd (every number is read from results/, seer/ and the repository files).

Software: R 4.6.1, PostgreSQL 17, dbt-postgres 1.12, Quarto 1.10, OHDSI ETL-Synthea 2.1, DataQualityDashboard 2.8.9. Library versions are recorded in results/sessionInfo.txt.

9 Use of AI tools

I used AI assistants (Claude) as tools in this project: to write and debug code, run checks, and draft documentation. The research question, study design, cohort and NSCLC definitions, SEER extraction, methodological choices and interpretation of the results are my own. I reviewed every output, verified results against their sources (including the SEER*Stat exports and the OMOP source data), and I am responsible for all content in this repository. Commits made with AI assistance carry Co-Authored-By trailers.

Code and results: https://github.com/erickyegon/oncology-rwe-nsclc. Medium article.