MLOps guide
What to know before you hire an MLOps services company
The questions a product or data lead should settle before paying anyone to run models in production: what MLOps covers, whether you need it yet, which platform, and what it costs to keep running.
What does an MLOps service actually include?
MLOps is the engineering that turns a trained model into a service you can update safely. In practice that means five pieces: a training pipeline that runs from code on a schedule or trigger; experiment tracking, so any result can be reproduced; a model registry that records which version is live; a deployment path with tests and a comparison against the current model; and monitoring that watches the model's inputs, outputs and accuracy.
What it does not include is building the model. If your team has no model yet, that is machine learning development work. MLOps starts once there is something worth deploying, and it pays for itself the second and third time that model is retrained.
When do you need MLOps, and when is it overkill?
You need it when a model affects revenue or risk and has to keep changing: fraud scores, prices, recommendations, demand forecasts. The signs you are late are familiar. Nobody can say which data trained the live model, retraining means a person re-running a notebook, or you learned about a bad model from a customer.
It is overkill for a single model scored once a month by a stable process. A scheduled job, a saved model file with a version number and a monthly accuracy check may be enough, and we will say so rather than sell a platform. The right amount of MLOps grows with the number of models and how often they change.
SageMaker, Vertex AI, Azure ML or open-source tools - which should you use?
Start with the cloud your data already lives in. Moving training data between clouds costs money and adds a security review, so a team on AWS usually lands on SageMaker, a Google Cloud team on Vertex AI and an Azure team on Azure Machine Learning. The managed platforms spare you from running servers for pipelines, registries and endpoints.
Open tools - MLflow for tracking and the registry, Kubeflow or Airflow for pipelines, containers on Kubernetes for serving - make sense when you run on more than one cloud, must run on-premises, or want to avoid writing every pipeline against one vendor's SDK. Mixing is common, for example MLflow tracking on top of a managed training service.
How do drift monitoring and retraining triggers work?
Two kinds of drift matter. Data drift means live inputs no longer look like the training data: a new customer segment, a changed upstream field, a seasonal swing. Prediction drift means the outputs have moved, such as the share of transactions flagged as fraud doubling overnight. Both can be measured the moment predictions are made. Accuracy itself can only be measured once true outcomes arrive, which may take days or weeks.
A retrain can be fired by any of the triggers below. Whatever fires it, the new model goes live only if it beats the current one on recent data.
- A schedule, for data that moves at a known pace.
- A drift threshold on key features or on the prediction mix.
- A batch of new labels large enough to test on.
- Measured accuracy falling below an agreed floor.
What is LLMOps, and how is it different from MLOps?
LLMOps applies the same discipline to products built on large language models, but the moving parts differ. You rarely retrain the model. Instead you change prompts, retrieval settings and the model version you call, and each change can alter answers in ways that are hard to spot. So prompts are versioned like code, and every change runs against an evaluation set of real questions with expected answers before it ships.
Guardrails check inputs and outputs for leaked personal data, off-topic requests and unsafe replies. Token spend is tracked per feature and per customer, because a longer prompt or a switch to a larger model can multiply the bill. For retrieval-based products see RAG development; for the model work itself, LLM development.
What does it cost to run models in production?
Three lines make up the bill. Serving: a tabular model behind an API usually runs on ordinary CPU containers, while deep learning and LLM serving need GPUs, which cost far more per hour and should scale down when idle. Training: scheduled retrains are modest unless they need GPUs or very large data. Monitoring and storage: logs, prediction history and feature data grow every day unless retention limits are set.
The main cost levers are batch scoring wherever real time is not needed, right-sized instances, request batching and quantized models on GPUs, and scale-to-zero for bursty traffic. The recurring cost teams forget is people: someone has to answer alerts and approve retrains. Our retainer from $2,299/month covers that after the 60 days of included support.
How do you evaluate an MLOps vendor?
Ask them to show rather than describe. These five questions separate a working platform from a slide deck.
A vendor whose answers all depend on its own proprietary platform is selling lock-in along with the service. With Miracuves, the pipelines, configs and runbooks sit in your repositories and your cloud, and your own engineers can run them after handoff.
- Can they reproduce your current live model from raw data before changing anything?
- Will pipelines and infrastructure be written as code, in your repository and your cloud account?
- How exactly is a new model promoted, and what happens step by step during a rollback?
- What do they monitor besides CPU and memory - inputs, predictions, accuracy, cost?
- Who answers a drift alert after handoff, and what does the runbook tell them to do?