AI-300 cheat sheet
Last-minute reference for AI-300. The model lifecycle domain is the largest; the three generative AI domains together are worth 40–55%.
Exam facts
| Duration | 120 minutes |
| Passing score | 700 / 1000 |
| Price | Set per country or region; shown at booking |
| Language | English only |
| Replaces | DP-100 |
Skill areas
| Area | Weight |
|---|---|
| Implement machine learning model lifecycle and operations | 25–30% |
| Design and implement a GenAIOps infrastructure | 20–25% |
| Design and implement an MLOps infrastructure | 15–20% |
| Implement generative AI quality assurance and observability | 10–15% |
| Optimize generative AI systems and model performance | 10–15% |
Requirement to answer
| Requirement | Answer |
|---|---|
| Recreate a workspace identically per environment | Bicep + Azure CLI |
| No stored Azure secret in GitHub | OIDC federated credential |
| Approval before prod deployment | GitHub environment with required reviewers |
| Same model or environment in dev, test and prod | Registry |
| Read storage without keys | Identity-based datastore + managed identity |
| Training compute that costs nothing idle | Compute cluster, min nodes 0 |
| Track parameters and metrics with one line | mlflow.autolog() |
| Search a continuous range and stop bad runs | Random sampling + bandit policy |
| Explain predictions, find failing cohorts | Responsible AI dashboard |
| Real-time scoring | Managed online endpoint |
| Score millions of files overnight | Batch endpoint |
| Test on real traffic, no user impact | Mirrored traffic |
| Canary release and instant rollback | Traffic split between deployments |
| Alert when inputs change | Model monitor with data drift signal |
| App calls a model without a key | Managed identity + RBAC on Foundry |
| Steady high volume, predictable latency and cost | Provisioned throughput (PTUs) |
| Data must stay in a geography | Regional or data zone deployment, not global |
| Keep the tested model version | Pinned version, no auto-upgrade |
| Answer contains facts not in the context | Groundedness |
| Answer doesn’t address the question | Relevance |
| Harmful content before launch | Risk and safety evaluators, adversarial data |
| Why did this one request fail? | Tracing in Application Insights |
| Exact codes missed by vector search | Hybrid search |
| Consistent style or format | Supervised fine-tuning |
| Preferred vs rejected answer pairs | DPO |
Sweep sampling
| Sampling | Continuous values | Early termination |
|---|---|---|
| Grid | No | Yes |
| Random | Yes | Yes |
| Bayesian | Yes | No |
RAG levers
Chunk size and overlap, top-k, similarity threshold, hybrid search, semantic ranking, embedding model. Changing the embedding model means re-embedding everything.
Traps
- A traffic split returns the new model’s answers to users. Mirroring does not.
- Archiving a model version does not delete it.
- Disabling public access breaks things unless storage, registry and DNS are private too.
- Fluent and relevant is not grounded.
- Evaluations that score everything zero usually mean a wrong column mapping.
- Fine-tuning is not for facts that change. Use RAG.
- Testing on synthetic data from the same run overstates performance.
- DP-100 material covers the ML half and misses GenAIOps and IaC.
Night-before checklist
- Name the five areas and which is largest
- Online versus batch endpoint, traffic split versus mirroring, cold
- The four quality metrics and what each measures
- Check ID and proctoring rules; see exam day
Take the 20-question practice test and check your weakest area.