AI Infrastructure & Reliability Engineer (Staff / Lead)
The engineer who makes a small, senior team run faster, steadier, and leaner.
About this role
Mlytics is transforming from a multi-CDN infrastructure company into an AI Decision Intelligence company. Our product lines include an AI-answer distribution system (Cortex) and an intelligent routing platform. The engineering team is small and senior, and we use AI to amplify what each person can ship.
This is not a traditional platform role. We don’t build an internal developer platform for its own sake, and we don’t set up release-approval gates. Your role is an amplifier: build the reliability mechanisms, own AI cost, and create the engineering backbone that lets the team dare to touch legacy. The mechanisms you define become the company’s engineering standards — and the cost optimizations you make show up directly on the company’s P&L.
What you’ll do
- Design reliability standards and on-call mechanisms — define the reliability standards and escalation mechanisms for all production services (AI products and the CDN routing platform), so each team can operate autonomously while you own the mechanism and serve as the last line of defense.
- Build AI cost governance (FinOps) — break down AI/infra cost by model, tenant, and feature; build an attribution model and budget visibility; and lead cross-team cost reduction. This one flows straight to gross margin.
- Set platform architecture direction — judge the adopt-vs-build boundary, prefer mature tools and managed services, and let the platform support product growth without growing headcount.
- Institutionalize DevEx — design shared deployment, observability, and testing scaffolding (a golden path) that makes “daring to touch legacy” the team norm.
- Technical leadership — as the technical mentor for infra/reliability, set direction and turn key knowledge into mechanisms. We don’t accept any system living only in one person’s head.
- Product-side security execution — implement security controls inside product/cloud systems in line with security policy and threat models.
Who we’re looking for
- 8+ years in platform / SRE / cloud engineering, having led cross-team reliability or platform-evolution initiatives.
- You’ve designed an SLO regime, on-call/escalation mechanisms, or a golden path and successfully driven adoption.
- Hands-on AI/infra cost governance (FinOps, COGS breakdown, and cost reduction).
- Fluent with AWS/GCP and Kubernetes (EKS) production environments and multi-cloud operations.
- You can act as a technical mentor and set direction; you have a track record of safely evolving hard-to-test legacy systems.
Nice to have: LLM-serving operations experience (inference reliability, model routing, fallback); Databricks or large-scale data-pipeline operations; a history of replacing headcount growth with automation on a small team.
You might not be a fit if you
- Want to build your own platform from scratch (we adopt over build).
- Are used to being the pre-release approval gate (we want an amplifier, not a gatekeeper).
- Hold systems together with personal memory and heroic firefighting (we want mechanisms, not heroes).
Why now
- Your leverage is enormous. A small, senior team means the mechanisms you define become company-wide standards directly — no layers of approval.
- Cost = margin. The FinOps you do isn’t an internal report; it’s a number you can see on the P&L.
- AI-native ways of working. We use AI tools like Claude Code as part of daily development — your output gets amplified by AI.
How to apply
Send us something that shows how you think about reliability and cost — an SLO regime you designed, a golden path you drove to adoption, a cost-reduction story with real numbers behind it. We care about what you’ve shipped more than what tools you’ve used.
📧 [email protected]