Senior AI Platform Finops Engineer
·
· 7+ years of experience in production operations, SRE, platform support, DevOps, or cloud infrastructure engineering.
· Strong hands-on experience supporting distributed systems in production, including Linux systems, containers, CI/CD, and cloud-native services.
· Strong experience in monitoring, logging, alerting, incident response, and service health management.
· Experience writing and maintaining operational runbooks, troubleshooting guides, and recovery procedures.
· Experience with identity, access control, audit logging, secrets handling, and security-aware platform operations.
· Experience supporting AI, data, or analytics platforms in production environments is strongly preferred.
· Strong scripting or development capability in Python and other automation-friendly languages.
· Ability to collaborate effectively across engineering, security, support, and business teams during operational events.
Preferred qualifications
· Experience with enterprise AI platform operations, including model-serving services, agent runtimes, prompt governance, or evaluation systems.
· Experience with Databricks, Dataiku, MLflow, Airflow, or similar platforms requiring production-grade operational governance.
· Experience with AKS or Kubernetes-backed platform operations, including autoscaling, certificate management, identity delegation, and cluster troubleshooting.
· Experience with Apigee X or similar API management platforms for production operations and secure service exposure.