Senior AI Platform Ops Engineer
· 8+ years of experience in DevOps, platform engineering, cloud infrastructure, or site reliability engineering, with strong production operations ownership.
· Strong hands-on experience with cloud platforms, microservices, CI/CD, containers, automation, and infrastructure design.
· Strong experience with Docker, Kubernetes, Linux administration, Git workflows, and system administration fundamentals.
· Hands-on development experience in Python and at least one additional language such as Node.js or Java.
· Experience with MLOps concepts such as model packaging, artifact management, deployment pipelines, inference operations, and rollback strategies.
· Experience operating AI or data platforms such as Databricks, MLflow, Airflow, or similar enterprise ML tooling.
· Strong understanding of observability, logging, alerting, and performance monitoring for distributed systems and AI workloads.
· Experience with secrets management, access control, and auditability in enterprise environments.
· Working knowledge of Apigee X or similar enterprise API gateway platforms.
Preferred qualifications
· Experience with AI gateway architecture, model catalogs, agent platforms, or MCP hosting patterns.
· Experience with vector databases, graph databases, feature stores, and RAG pipeline operations.
· Experience with AKS or other managed Kubernetes services, managed identities, Key Vault, TLS, SSO, and enterprise platform runbooks.
· Experience working in Agile delivery models with strong cross-functional communication and a bias for action.