Role OverviewAs a Principal Infrastructure Engineer / Tech Lead, you'll be one of the foundational infrastructure engineers responsible for designing and scaling the cloud systems that power custom AI model training and inference. This is an opportunity for a hands-on technical leader who wants to own infrastructure architecture end-to-end while remaining deeply involved in engineering execution, cloud systems, ML infrastructure, reliability, security, and cost optimization.
What You Will Do
Own and scale the cloud infrastructure powering custom AI model training and inference for enterprise customers, serve as a technical lead for infrastructure and platform engineering initiatives, design reliable and scalable cloud architecture across production environments, and build and maintain deployment systems for AI models and production services.
Why It Might Be a Fit
This is an opportunity to become a foundational infrastructure leader at a highly capital-efficient AI company solving complex problems at the intersection of machine learning, cloud infrastructure, and enterprise technology. You'll own the infrastructure powering custom AI models for major enterprise customers rather than maintaining mature systems at a large organization.
Requirements
- 7+ years of experience in cloud infrastructure, DevOps, platform engineering, or related infrastructure roles
- Experience serving as a technical lead for an infrastructure or platform engineering team
- Strong hands-on individual contributor experience in a current or recent role
- Experience owning infrastructure projects from architecture through production
- Experience working at startups or early-stage technology companies
- Experience building and operating production cloud infrastructure
- Experience working with large-scale or highly reliable production systems
- Experience with infrastructure architecture and technical leadership
- Experience collaborating closely with engineering and technical leadership
- Experience supporting ML workflows or research infrastructure preferred
- Strong ability to operate independently in ambiguous startup environments
- Strong technical communication and problem-solving skills
- Experience mentoring other infrastructure or software engineers
- Strong cloud infrastructure experience
- Strong AWS experience preferred
- GCP or Azure infrastructure experience considered transferable
- Strong Python coding fluency
- Experience with Infrastructure-as-Code tools such as Terraform
- Strong CI/CD experience
- Experience designing and maintaining production deployment systems
- Experience with cloud networking and infrastructure architecture
- Experience with production monitoring, logging, and observability
- Experience with infrastructure reliability and incident response
- Experience managing cloud infrastructure costs
- Experience with containerized production workloads
- Experience with scalable backend or platform infrastructure
- Strong understanding of distributed systems fundamentals
- Experience supporting high-compute workloads preferred
- Experience scaling GPU workloads preferred
- Experience with MLOps or ML infrastructure preferred
- Experience with ETL and data pipelines preferred
- Experience with model deployment and lifecycle management preferred
- Experience with infrastructure security and data isolation
- Experience with SOC2 or comparable security compliance initiatives preferred
- Strong systems design and architecture skills
- Ability to troubleshoot complex infrastructure problems
- Ability to balance reliability, scalability, performance, and cost
- Ability to remain hands-on while operating as a technical lead
- Bachelor's degree in Computer Science, Engineering, or related technical field preferred
- Strong computer science fundamentals
- Equivalent practical infrastructure engineering experience accepted
Benefits
- Base Salary: $200,000 – $250,000
- Potential flexibility above $250K for exceptional candidates
- Competitive startup equity targeted around the 90th percentile relative to market for strong candidates
- Opportunity to work as one of the earliest infrastructure hires
- High ownership over foundational cloud and ML infrastructure
- Direct influence over infrastructure architecture and technical strategy
- Opportunity to build systems from the ground up
- Exposure to GPU-heavy AI infrastructure and custom ML model deployment
- Opportunity to solve complex infrastructure challenges at enterprise scale
- Direct collaboration with AI researchers and technical leadership
- Opportunity to mentor and shape the infrastructure engineering function
- High autonomy in a small, technically strong engineering organization
- Relocation assistance to the San Francisco Bay Area
- H-1B and OPT transfer support
]]>