Senior AI Platform Finops Engineer
·
- 7+ years of experience in production operations, SRE, platform support, DevOps, or cloud infrastructure engineering.
- Strong hands-on experience supporting distributed systems in production, including Linux systems, containers, CI/CD, and cloud-native services.
- Strong experience in monitoring, logging, alerting, incident response, and service health management.
- Experience writing and maintaining operational runbooks, troubleshooting guides, and recovery procedures.
- Experience with identity, access control, audit logging, secrets handling, and security-aware platform operations.
- Experience supporting AI, data, or analytics platforms in production environments is strongly preferred.
- Strong scripting or development capability in Python and other automation-friendly languages.
- Ability to collaborate effectively across engineering, security, support, and business teams during operational events.
Preferred qualifications
- Experience with enterprise AI platform operations, including model-serving services, agent runtimes, prompt governance, or evaluation systems.
- Experience with Databricks, Dataiku, MLflow, Airflow, or similar platforms requiring production-grade operational governance.
- Experience with AKS or Kubernetes-backed platform operations, including autoscaling, certificate management, identity delegation, and cluster troubleshooting.
- Experience with Apigee X or similar API management platforms for production operations and secure service exposure.