Research capability

De-identification and clinical data

Retrospective clinical studies from protocol and cohort construction through analysis, reporting, de-identification, and defensible data release.

We run retrospective clinical data programs from the raw export through analysis and reporting. When data or derived products need to leave the institution, we extend that work through de-identification and release documentation.

Retrospective study design

We translate the scientific question into inclusion and exclusion criteria, index dates, exposure and comparator definitions, outcomes, covariates, and a statistical analysis plan. Feasibility checks happen early, before a protocol depends on variables the source system does not capture reliably.

The same plan defines sensitivity analyses and negative controls where they can expose an alternative explanation. Predictive and causal questions are kept separate because they require different evidence.

Building research databases from raw records

Before anything else can happen the data has to be assembled. Hospital exports come from systems built for billing and bedside care: overlapping timestamps, units that change halfway through a decade, three different codes for the same lab test.

We do the unglamorous part of that work.

  • Retrospective extraction from EHR and adjacent systems, with the extraction logic version-controlled rather than run by hand
  • Integration of vital signs, laboratory results, medications, procedures, and clinical notes into one patient timeline
  • Temporal alignment, so an event that happened at 03:00 in one source is not five hours off in another
  • Quality checks that catch what a schema validator will not, such as a ward whose blood pressures stop being recorded for six weeks

De-identification

Tabular records require structured guarantees removing direct and indirect identifiers. Consistent date shifting supports chronological research. Fields which are free-text require a thoughtful decision between omission, allow-listing, or applying machine learning techniques. We have extensive experience applying state of the art de-identification approaches to a wide variety of data types.

Safe Harbor, Expert Determination, and the paperwork

HIPAA Safe Harbor has become somewhat of a standard across the field, though its implementation requires careful thought. Expert Determination is generally preferred for organizations interested in explicitly modeling the risk, and potentially broadening the data elements included in the resultant dataset. We work through both routes.

Clinical evidence generation

Analysis of clinical data requires careful steps made with a full understanding of the data generation process. Missing lab values are informative; models can be tricked by shortcut features; and so on. We develop and validate phenotypes, handle missingness and censoring, fit statistical or machine-learning models, and test whether conclusions survive alternative cohort and endpoint definitions. Results are reported with the protocol, code, cohort flow, data-quality findings, and limitations needed for scientific review.

Want to talk it through?

Send us the shape of the problem. We will say quickly whether we are the right people for it.

Email us