Calister Nnona

Research

Papers and technical reports, each with its code and, where one exists, the full technical report behind it.

Preprints

CoRTeC: Cohort-Conditioned Differentially Private Synthetic Tabular Data from a Frozen Language Model

Preprint, 2026

Institutions in regulated domains want to train on, and share, records they cannot release. CoRTeC spends the privacy budget once, on a differentially private release of cohort-conditioned statistics, and lets a frozen, un-finetuned language model decode it into records; because the generator never sees a private row, unlimited records cost no further budget. At ε = 2 on three regulated-domain datasets the output is within 0.016 of the most accurate marginal-based method on fidelity, and tree models trained on it match models trained on a real sample. Along the way the paper shows that the two acceptance criteria the industry uses for synthetic data are saturated: real data and a naive baseline score the same on them, so they cannot tell good synthetic data from bad.