Open role · Remote (US)

Senior Data / ML Engineer

$150,000 – $170,000 USD base salary

We're looking for someone who can co-own how we decide data is safe to release, and do the engineering that makes that decision real at customer scale. You'll be the third technical person on the team.

About the role

Much of the world's most valuable data — healthcare records, financial transactions, clinical trial data — can't be easily analyzed because it contains sensitive personal information. Today, access to this data is controlled through slow, manual governance processes that can take months to approve.

There are only two honest ways to make that data usable. You can generate a synthetic version that's safe to share. Or you can compute on the real thing and let only safe answers back out. We build both — and underneath, they're the same question: what could someone infer about a real person from what you just released?

Answering that question is the job. It doesn't have a settled answer, a benchmark, or a textbook. It also isn't a research position: every answer has to become working code running against real customer data, and then has to hold up when a hospital's privacy officer goes looking for a reason to say no.

What you'll work on

The role has two centers of gravity, split roughly evenly.

Privacy and disclosure methodology. On the synthetic side, we've designed a state-of-the-art privacy methodology working with third-party auditors that allows us to generate de-identified data at scale. On the real-data side, we're doing the same to determine which aggregates can be released exactly, which have to be coarsened, and which get suppressed entirely.

This is more than an implementation exercise - we're building on existing privacy frameworks and combining those abstractions with systems concepts to make them easy to use in some of the largest health systems in the country. Most of our stakeholders aren't privacy experts, so we need to be able to describe these concepts in ways that make the "so what" clear to a wide range of experts.

The pipeline that produces all of it. Our synthesis pipeline is a typed DAG on Ray: cohort selection, reshaping arbitrary customer schemas into a canonical form, encoding records, distributed training of a database-level generative model, generation, and evaluation. You'd own making it correct and fast against real customer databases — which includes the unglamorous half, where a customer's data synthesizes badly and someone has to figure out why.

Our tech stack:

How we work

We're intentionally keeping the team small while we build the core product. There are two technical contributors today; you'll be the third. You'll help set the direction for the product and your work will be a meaningful fraction of everything we ship.

We're a fully remote company distributed across the US, and everybody can work from anywhere.

Who thrives here

You'll succeed here if you enjoy:

You may struggle here if you prefer:

Baseline qualifications:

Things that help, none of which are required:

If you're excited about the problem but unsure whether you meet every qualification, we'd still like to hear from you. We care much more about ownership and technical judgment than any specific resume.