Role OverviewYou will own the Azure platform behind SimpliGov's AI-native delivery model—infrastructure, Kubernetes, networking, observability, AI serving, and cost discipline. This is hands-on production engineering within a FedRAMP-conscious environment, where security, auditability, and reliability are core responsibilities.
What You Will Do
Deploy and operate our Azure platform: AKS, networking, identity, storage, and environments from development through production. Own infrastructure as code end to end: environments are reproducible, drift is detected, and nothing reaches an environment without platform visibility.
Why It Might Be a Fit
You will partner with AI Operations on the deploy-and-release path: Octopus Deploy, environment promotion, progressive rollout, and rollback. Build platform reliability: monitoring, alerting, incident response, and capacity planning.
Requirements
- 5+ years in DevOps, platform engineering, or site reliability engineering in SaaS environments
- Deep Azure experience: AKS, networking, identity (Entra), and monitoring; you have run production Kubernetes
- Infrastructure as code as your default (Terraform, Bicep, or similar), plus strong scripting; you automate before you document
- MLOps experience: deploying and operating LLM or ML systems in production, including model gateways, inference infrastructure, or AI observability stacks
- Demonstrated cost work: you can point to cloud spend you found, explained, and reduced
- Experience in compliance-heavy environments (FedRAMP, StateRAMP, SOC 2, or similar) is a strong plus
Benefits
- Medical, dental, and vision insurance plans, with significant employer contributions for employees AND dependents
- Company-sponsored life/disabilities insurances
- 11 Paid holidays
- Flexible time off
- 401k plan with 4% employer match
- Monthly stipends for wellness and home office expenses
]]>