AI-300 ML model lifecycle explained

Updated September 28, 2026

Implement machine learning model lifecycle and operations is the largest AI-300 domain at 25–30%. It follows one model from training through registration and deployment to monitoring in production. Much of it will be familiar from DP-100, but the emphasis is on the operational end: how a model is versioned, rolled out safely, rolled back, and retrained when the data changes.

Orchestrating training

Jobs and MLflow

A command job runs a script on a compute target in an environment. Azure Machine Learning workspaces are MLflow tracking servers, so mlflow.autolog() or explicit mlflow.log_metric calls record parameters, metrics and artifacts without extra configuration inside a job. Comparing runs in the studio, or querying them with the MLflow client, is how you compare model performance across jobs.

AutoML and notebooks

Automated machine learning tries algorithms and featurisation for classification, regression, forecasting and some vision and NLP tasks, and ranks them by a primary metric you choose. Notebooks remain the place for exploration. The exam expects you to know when each fits, not to prefer one.

Hyperparameter tuning

A sweep job wraps a training job with a search space, a sampling algorithm and an early termination policy.

SamplingBehaviour
GridEvery combination; only discrete values
RandomRandom combinations; supports continuous ranges
BayesianPicks new values based on earlier results; does not support early termination

Early termination policies such as bandit, median stopping and truncation selection stop poor runs early to save compute.

Distributed training and pipelines

Large and deep learning models are trained across several nodes by setting a distribution type (for example PyTorch or MPI) and an instance count on the job. Pipelines chain components, pass outputs to inputs, reuse unchanged steps, and can be scheduled or triggered from a workflow.

Registration and versioning

Registering an MLflow model stores the artifacts with an MLmodel file describing its signature and flavour, which makes no-code deployment possible. Each registration creates a new version; tags and descriptions record lineage. Archiving hides an old version from default lists without deleting it, so existing references keep working.

When the model depends on features from a feature store, a feature retrieval specification is packaged with the model artifact so that inference looks up exactly the features the model was trained on.

Before promotion, a Responsible AI dashboard evaluates the model: error analysis, fairness, model interpretability and counterfactuals. It is the answer when a scenario asks how to find which data segments the model fails on, or which features drive its predictions.

Deploying to production

EndpointUse it when
Managed online endpointLow-latency, real-time scoring over HTTPS
Batch endpointScoring large volumes of files asynchronously

An online endpoint holds one or more deployments. Traffic is split between them by percentage, which gives you blue/green and canary releases: send 10% to the new deployment, watch its metrics, then move to 100% or back to 0%. Mirrored traffic copies a share of requests to a new deployment without returning its responses to clients, so you can test under real load with no user impact.

Troubleshooting starts with deployment logs and, for custom containers, the container logs. Local deployment with a local endpoint helps catch scoring script errors before paying for compute. Common failures: a scoring script whose init() cannot load the model, a missing package in the environment, or too little memory on the instance type.

Monitoring and maintenance

Model monitoring compares production data with a reference data set, usually the training data, and computes signals:

  • Data drift: input feature distributions have shifted.
  • Prediction drift: the distribution of outputs has shifted.
  • Data quality: nulls, type errors, out-of-range values.
  • Feature attribution drift: the features that drive predictions have changed.

Each signal has thresholds. Crossing one raises an alert, which can drive a retraining pipeline through an event or workflow rather than someone noticing a dashboard.

Sample questions

Question 1. You run a hyperparameter sweep over a continuous learning-rate range and want to stop runs that are clearly underperforming, to save compute. Which configuration fits?

  • A. Grid sampling with a bandit policy
  • B. Random sampling with a bandit policy
  • C. Bayesian sampling with a median stopping policy
  • D. An AutoML job with the learning rate as primary metric
Show answer

Answer: B

Random sampling supports continuous ranges and can be combined with an early termination policy such as bandit. Grid sampling only supports discrete choices. Bayesian sampling does not support early termination policies. AutoML does not sweep a custom training script’s parameters.

Want more questions like this? Full AI-300 practice tests →

Question 2. Before shifting any user traffic to a new model deployment on a managed online endpoint, you want to observe its latency and errors on a copy of real production requests. Clients must only ever receive responses from the current deployment. What should you use?

  • A. A 10% live traffic split to the new deployment
  • B. A batch endpoint that scores yesterday’s requests
  • C. Mirrored traffic to the new deployment
  • D. A local endpoint on a developer machine
Show answer

Answer: C

Mirrored traffic sends a copy of a percentage of requests to the new deployment and discards its responses, so it is tested under real load with no effect on clients. A 10% traffic split returns the new deployment’s responses to real users. A batch endpoint and a local endpoint do not see live requests.

Want more questions like this? Full AI-300 practice tests →

Question 3. A model in production shows no errors, but its business performance is slowly declining. Customer behaviour has changed since training. You want to be alerted automatically when the inputs no longer resemble the training data. What should you configure?

  • A. Model monitoring with a data drift signal against the training data and an alert threshold
  • B. An Azure Monitor alert on endpoint CPU utilisation
  • C. A Responsible AI dashboard generated at registration
  • D. A schedule that retrains the model every night
Show answer

Answer: A

A model monitor with a data drift signal compares production inputs with the training data as reference and raises an alert when a threshold is crossed. Endpoint CPU and error metrics would not reveal a change in data distribution. A Responsible AI dashboard is a point-in-time analysis. Retraining nightly treats the symptom without detecting it and wastes compute.

Want more questions like this? Full AI-300 practice tests →

What to practise

Train a model with MLflow autologging, sweep two parameters, register the best run as an MLflow model, and deploy it to a managed online endpoint. Add a second deployment, mirror traffic to it, then split traffic 90/10 and roll back. Finish by configuring a data drift monitor. Almost every objective in the domain is in that one exercise.