Open Label

Industry · 2026-07-18

What AI actually does in statistical programming right now

Judge by the conference agenda and the takeover is already here. PharmaSUG and PHUSE are full of sessions on large language models writing SDTM mappings, ADaM specs, and table code. Judge by the papers inside those sessions and the story gets quieter, more careful, and a lot more interesting for anyone who does this work for a living.

What is actually being built

Two 2025 papers, written independently by people who ship datasets for a living, converge on the same architecture. A senior statistical programming manager at Pacira described a GPT-based SDTM pipeline: study specs distilled into machine-readable metadata, templated prompts per domain, generated Python that builds DM or EX, then automated quality checks. When a check fails, the error message goes back to the model with context and it proposes a fix. A PHUSE paper from Sycamore Informatics describes the same shape for TLF code in SAS, with one blunt addition: "relying on a general-purpose LLM is insufficient and introduces unacceptable regulatory risk." Their design keeps a fine-tuned model inside validation gates, with the programmer as the final gateway.

Neither paper describes a programmer being replaced. Both describe the job moving up a level: from writing the derivation to specifying it precisely enough that a model can write it, then owning the QC that proves the output right. The spec was always the hard part. Now it is the product.

The missing numbers

Here is the tell worth sitting with. Neither practitioner paper publishes an accuracy rate, a time saving, or a failure count. Not one. The Pacira paper reports feasibility and lists its own failure modes: ambiguous specs produce wrong code, spec formats vary too much to standardize, complex errors still need a human.

The numbers live somewhere else. Certara says its CoAuthor software cuts regulatory first-draft time by 30 percent and that its new TFL Studio builds tables up to 50 percent faster. Training-industry blogs claim 95 percent SDTM automation and thousands of hours saved, with no named study behind the figures. None of these come with a sample size, a baseline, or a method. On our salaries page a number without an n does not get published. The same rule applies here: treat every unattributed percentage as marketing until someone shows the study.

Regulators moved before the tools did

The FDA issued its first draft guidance on AI in drug development in January 2025. It proposes a risk-based credibility framework: define the question, define the context of use, assess model risk, then prove credibility for that context. The EMA published its reflection paper in September 2024 with the same posture. Neither document bans AI-written code. Both put the burden of proof on the sponsor, which in practice means on the biometrics team signing the submission.

There is a wrinkle programmers should know. The FDA draft scopes itself to AI that produces information supporting regulatory decisions on safety, efficacy, or quality. Code generation sits in a gray zone: the model writes the program, but the program is deterministic and testable by every method the industry already trusts. That is exactly why the practitioner architectures lean so hard on validation gates. The compliance story runs through the QC evidence, not the model.

What it means for the job

No CRO or sponsor has announced statistical programmer layoffs attributed to AI, and we could not verify any. The pressures that are documented look familiar: offshoring, the slow shift from SAS toward R and Python, and flat headcount stretched across more studies. AI enters as an accelerant on those trends, not a separate wave. The skills the practitioner papers reward are metadata curation, spec discipline, and validation ownership, which is to say the senior half of the job description.

Government filings on our salaries page show what the ladder paid before any of this landed: senior statistical programmers at IQVIA filed at a median of $122,400, principal level at $150,565. Whether those numbers bend over the next two years is an empirical question, and self-reports will show it before any press release does. If you program for a living, your datapoint is the measurement.

Discussion

0 commentsAnonymous, verified members. House rules apply.

Nobody has weighed in yet. If this piece matches or misses your experience, say so below; one sentence is enough.

Your reply is kept while you verify your work email, then it lands here under your pseudonym.

Open Label is building the salary dataset this industry never had. Add your anonymous datapoint. Three minutes, no name, no email.