DigeHealth logo

Senior Machine Learning Engineer

DigeHealth

On-siteSan Francisco Bay AreaseniorPosted 2h ago

Job description

About the role

DigeHealth is building the revolution of gut health — a Swiss wearable startup creating the first continuous, non-invasive platform for gastro-intestinal health, on a dataset that exists nowhere else in the world, built hand in hand with leading Swiss university hospitals and annotated by practising gastroenterologists.

Reporting to the CTO, you will help build the next foundational model for gut activity: a self-supervised, multimodal architecture trained on continuous acoustic, physiological and contextual streams, learning general-purpose representations of gastro-intestinal dynamics that transfer across downstream biomarker tasks and improve as the data grows. Our next frontier is scaling from expertly labelled data to the far larger volume of unlabelled signal we are collecting. That is the heart of this role.

What you’ll do

  • Design, train and evaluate models on multimodal time series: acoustic, PPG and IMU.

  • Push our self-supervised and semi-supervised work forward, scaling from labelled to unlabelled data.

  • Work directly with the raw signal, from denoising and noise cancellation to feature representation.

  • Optimize models for edge deployment under real memory, latency and power constraints, keeping computation on device.

  • Build the training and evaluation infrastructure: reproducible experiments, versioned datasets, clinically meaningful metrics.

  • Take models to production, including serving, monitoring, drift detection and retraining.

  • Work with our clinical partners to define what the model should predict and how performance should be judged.

  • Contribute to publications and to our patent portfolio.

What we’re looking for

  • 5+ years building machine learning systems in production, not research alone.

  • Deep expertise in time series, audio or biosignal modelling.

  • Strong Python and fluency with PyTorch or equivalent.

  • You can reason about why a model works, diagnose failure modes and design experiments that answer questions.

  • Intellectual honesty about performance, especially where clinical relevance is at stake.

  • Comfortable with messy real-world data and a problem with no benchmark to beat — you care about work that holds up under clinical scrutiny, not only a favourable test split.

  • Clear communication with clinicians and non-technical colleagues.

Nice to have

  • Self-supervised or semi-supervised learning on large unlabelled datasets.

  • Edge or embedded deployment, including quantization and pruning.

  • Biosignals, medical devices or another domain where model errors carry real consequences.

  • A publication record.

What we offer

Access to a dataset that does not exist anywhere else and the freedom to define how it is used, meaningful equity, direct collaboration with the founding team and our clinical partners, and support for publishing your work.