Member of Technical Staff, Data Engineering

Mithrl
San Francisco, CA
Job Description
Role Overview

The role is to build and own an AI-powered ingestion & normalization pipeline to import data from various sources, develop robust schema mapping, coercion, and conversion logic, and ensure all transformations execute during ingestion.

What You Will Do

Build and own an AI-powered ingestion & normalization pipeline, develop robust schema mapping, coercion, and conversion logic, use LLM-driven and classical data-engineering tools to structure messy tabular data, and build validation, verification, and quality-control layers.

Why It Might Be a Fit

The role requires strong fluency in Python, data processing tools, and experience dealing with messy Excel / CSV / spreadsheet-style data, as well as the ability to combine classical data engineering with LLM-powered data normalization / metadata extraction / cleaning.

Requirements

  • 5+ years of experience in data engineering / data wrangling with real-world tabular or semi-structured data
  • Strong fluency in Python, and data processing tools (Pandas, Polars, PyArrow, or similar)
  • Excellent experience dealing with messy Excel / CSV / spreadsheet-style data — inconsistent headers, multiple sheets, mixed formats, free-text fields — and normalizing it into clean structures
  • Comfort designing and maintaining robust ETL/ELT pipelines, ideally for scientific or lab-derived data
  • Ability to combine classical data engineering with LLM-powered data normalization / metadata extraction / cleaning
  • Good communication skills; able to collaborate across teams (product, bioinformatics, infra) and translate real-world messy data problems into robust engineering solutions

Benefits

  • Comprehensive PPO health coverage through Anthem (medical, dental, and vision)
  • 401(k) with top-tier plans
]]>