Open role · Remote (US)
Senior Data / ML Engineer
We're looking for someone who can co-own how we decide data is safe to release, and do the engineering that makes that decision real at customer scale. You'll be the third technical person on the team.
About the role
Much of the world's most valuable data — healthcare records, financial transactions, clinical trial data — can't be easily analyzed because it contains sensitive personal information. Today, access to this data is controlled through slow, manual governance processes that can take months to approve.
There are only two honest ways to make that data usable. You can generate a synthetic version that's safe to share. Or you can compute on the real thing and let only safe answers back out. We build both — and underneath, they're the same question: what could someone infer about a real person from what you just released?
Answering that question is the job. It doesn't have a settled answer, a benchmark, or a textbook. It also isn't a research position: every answer has to become working code running against real customer data, and then has to hold up when a hospital's privacy officer goes looking for a reason to say no.
What you'll work on
The role has two centers of gravity, split roughly evenly.
Privacy and disclosure methodology. On the synthetic side, we've designed a state-of-the-art privacy methodology working with third-party auditors that allows us to generate de-identified data at scale. On the real-data side, we're doing the same to determine which aggregates can be released exactly, which have to be coarsened, and which get suppressed entirely.
This is more than an implementation exercise - we're building on existing privacy frameworks and combining those abstractions with systems concepts to make them easy to use in some of the largest health systems in the country. Most of our stakeholders aren't privacy experts, so we need to be able to describe these concepts in ways that make the "so what" clear to a wide range of experts.
The pipeline that produces all of it. Our synthesis pipeline is a typed DAG on Ray: cohort selection, reshaping arbitrary customer schemas into a canonical form, encoding records, distributed training of a database-level generative model, generation, and evaluation. You'd own making it correct and fast against real customer databases — which includes the unglamorous half, where a customer's data synthesizes badly and someone has to figure out why.
Our tech stack:
- Python, Ray, and PyTorch for modeling and data pipelines
- Node for backend services, TypeScript and React for web UI
- Kubernetes on all major clouds
How we work
We're intentionally keeping the team small while we build the core product. There are two technical contributors today; you'll be the third. You'll help set the direction for the product and your work will be a meaningful fraction of everything we ship.
- Small team, high ownership
- Engineers talk directly with customers — you'll debug a dataset on a call and explain privacy results to compliance teams
- We optimize for shipping and learning quickly
- Technical decisions are made by the people closest to the problem
We're a fully remote company distributed across the US, and everybody can work from anywhere.
Who thrives here
You'll succeed here if you enjoy:
- Identifying problems and carrying the solution all the way through launch
- Owning problems end-to-end, from the statistical proof-of-concept to the code running in production
- Defending a methodological choice to someone who doesn't want to accept it
- Moving between a modeling question and a cluster failure in the same afternoon
You may struggle here if you prefer:
- Fully defined requirements and a well-groomed backlog
- Doing the modeling and handing implementation to someone else, or working within a single layer of the product stack
- Research output that stops at a finding rather than a shipped feature
Baseline qualifications:
- ~5+ years of applied ML or data science experience inside a product or data infrastructure company
- Strong Python, and a willingness to work outside it
- Experience writing production-quality code
- Comfort working directly with customers
Things that help, none of which are required:
- Experience with healthcare or other regulated data (EHR, claims, OMOP, HIPAA de-identification)
- Distributed compute at scale, particularly Ray or Spark
- Generative modeling depth, especially sequence models and the evaluation of generated output
If you're excited about the problem but unsure whether you meet every qualification, we'd still like to hear from you. We care much more about ownership and technical judgment than any specific resume.