I build machine learning systems for teams working with sensitive, regulated data. Most recently, generative models that let a clinical research team experiment on patient-like data without ever touching a real patient record.
Diffusion models that generate synthetic tabular clinical data, preserving the statistical structure of real datasets while keeping protected health information out of the pipeline entirely.
LLM embedding and clustering pipelines over medical literature, turning a pile of search results into findings a research team could actually act on.
Prompt-engineered documentation workflows (built on Claude) that cut the manual-writeup time clinical researchers were spending on study documentation.
Generating synthetic data that looks statistically right is the easier half of this work. Getting clinical researchers, people trained to be skeptical of anything that isn't a real measurement, to actually rely on it downstream is the harder one. On the engagement below, that meant running statistical comparisons between synthetic and real distributions, walking non-technical stakeholders through the results directly, and feeding their pushback back into the validation checks until the pipeline was something people trusted enough to use.
This sits on top of 10+ years of production ML work (deep learning, computer vision, LLMs) across healthcare, finance, retail, and security, most recently as a Senior Machine Learning Engineer at Tecla/Incyte. It's also not a new interest: my MSc thesis (Universidad de Chile, 2016) was on automated blood vessel segmentation in retinal images for diabetic retinopathy and cardiovascular screening, an applied medical-imaging problem I worked on years before generative AI made “healthcare AI” a category. That work is documented here too.
If your team is dealing with data access, privacy, or trust problems around sensitive data, I'd like to hear about it.
sebastian.cepeda.fuentealba@gmail.com