← Back to job listings
VT
MLOps Engineer / ML Platform Engineer
Valce Talent Solutions · Mexico
About The Role
AI Engineer – AI Ops (Monitoring AI Systems in Production)
Role purpose
An AI Engineer in AI Ops & Governance is responsible for operating, monitoring, and maintaining AI models and agent-based systems once they are live in production. The role ensures AI systems remain reliable, performant, compliant, and aligned with business expectations over time, closing the “last‑mile” gap between model development and sustainable production use
Key Responsibilities
Production Monitoring & Health
- Monitor AI models and agents in production for performance, latency, errors, and availability.
- Track statistical health indicators such as model drift, data distribution changes, and output stability.
- Observe business KPIs linked to AI behaviour (e.g. accuracy impact, false‑positive cost, efficiency).
Incident Management & Recovery
- Detect and triage production incidents related to AI behaviour or degradation.
- Execute rollbacks, throttling, or model disabling where thresholds are breached.
- Support root‑cause analysis and post‑incident reviews to prevent recurrence.
Model & Agent Lifecycle Operations
- Support deployment, versioning, and release of AI models and agents using CI/CD‑style pipelines.
- Maintain registries and metadata covering model ownership, lineage, risk classification, and approvals.
- Support models and agents move safely through environments (dev → test → production).
Governance, Risk & Compliance
- Ensure AI systems adhere to Responsible AI principles, internal controls, and audit requirements.
- Maintain audit trails, logs, and approval artefacts required by risk, compliance, and regulators.
- Support fairness, bias, explainability, and transparency monitoring in production.
Tooling & Platform Integration
- Integrate AI systems with monitoring, logging, and alerting platforms (e.g. dashboards, metrics stores).
- Work with cloud infrastructure (containers, event streaming, APIs) supporting scalable AI operations.
- Collaborate with product, engineering, and data teams to standardise AI Ops patterns and blueprints.
Skills & Experience (Baseline)
- Strong Python skills and experience supporting ML or LLM‑based systems.
- Understanding of Model Ops / MLOps, especially the operational phase after deployment.
Experience with
- Monitoring and logging systems
- CI/CD pipelines
- Containerised deployments (e.g. Docker‑based runtimes)
- Familiarity with cloud platforms (Azure preferred) and production troubleshooting.
- Ability to work cross‑functionally with product, data science, engineering, and risk teams
Similar roles you might like
See all →1S
ETL Software Engineer
120 SFDC Mexico S. de R.L. de C.V.
Salary not disclosedPosted today
NY
SFMC Email Tech Developer
Nir Yu
Salary not disclosedPosted today
NY
Fullstack Developer (PHP React)
Nir Yu
Salary not disclosedPosted today
VT
DevOps/SRE Engineer
Valce Talent Solutions
Salary not disclosedPosted today
MC
Product Engineer
MEX Centro Tecnico Herramental S. de R.L. de C.V.
Salary not disclosedPosted today
C
Fullstack AI Enginner
Creai
Salary not disclosedPosted today
BM
Desarrollador LRBA Junior
BABEL México
Salary not disclosedPosted today
BM
Desarrollador webMethods Semi Senior
BABEL México
Salary not disclosedPosted today
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
