Free AI-300 practice test: 20 questions

Updated September 28, 2026

Twenty questions across the five AI-300 skill areas, weighted roughly as the real exam is: four on MLOps infrastructure, six on the model lifecycle, four on GenAIOps infrastructure, three on quality and observability and three on optimisation. These are original scenario questions. Read what the requirement actually asks before picking an answer.

Design and implement an MLOps infrastructure

Question 1. Training jobs need a machine that scales to zero nodes when no jobs are queued and up to eight nodes under load. Which compute target should you create?

  • A. A compute instance
  • B. A compute cluster with minimum nodes 0 and maximum nodes 8
  • C. An attached Kubernetes cluster with one node
  • D. A managed online endpoint
Show answer

Answer: B

A compute cluster scales between a minimum and maximum node count, and a minimum of zero means no cost while idle. A compute instance is a single development machine that bills until stopped. Attached Kubernetes compute requires a cluster you run yourself, and a managed online endpoint is for inference, not training.

Want more questions like this? Full AI-300 practice tests →

Question 2. A team wants to recreate its Azure Machine Learning workspace, storage, Key Vault and Application Insights identically in three environments from its Git repository. What should it use?

  • A. A Bicep template deployed with az deployment group create
  • B. Step-by-step instructions for creating the resources in the portal
  • C. An Azure Machine Learning registry
  • D. A notebook that documents each resource’s settings
Show answer

Answer: A

A Bicep template deployed with the Azure CLI declares the workspace and its dependent resources as code, so every environment is created the same way from the repository. Portal clicks and exported screenshots cannot be repeated reliably, and a registry shares assets, not infrastructure.

Want more questions like this? Full AI-300 practice tests →

Question 3. A workflow must run the same pipeline YAML against a dev workspace on every pull request and against the prod workspace only after a named approver agrees. How should you configure GitHub Actions?

  • A. Store the prod pipeline in a separate repository
  • B. Schedule the prod job with a cron trigger at night
  • C. Add a comment in the workflow file asking people to check before merging
  • D. Define dev and prod GitHub environments and require a reviewer on prod
Show answer

Answer: D

GitHub environments with required reviewers pause a job until an approver agrees, and each environment can carry its own federated credential and variables for its workspace. A separate repository duplicates code. A cron schedule does not add approval, and a comment in the YAML enforces nothing.

Want more questions like this? Full AI-300 practice tests →

Question 4. A data scientist should be able to submit jobs and register models in a workspace, but must not be able to change the workspace's network or compute quota settings. Which approach fits best?

  • A. Assign the Owner role on the workspace
  • B. Assign the Contributor role on the resource group
  • C. Assign the AzureML Data Scientist role on the workspace
  • D. Assign the Reader role on the workspace
Show answer

Answer: C

The AzureML Data Scientist built-in role allows working with jobs, assets and models without permission to manage the workspace itself. Owner and Contributor both allow changing workspace settings. Reader cannot submit jobs.

Want more questions like this? Full AI-300 practice tests →

Implement machine learning model lifecycle and operations

Question 5. Training scripts run as command jobs in Azure Machine Learning. You want parameters, metrics and the model artifact recorded for every run with as little code change as possible. What should you add?

  • A. Call mlflow.autolog() in the training script
  • B. Print metrics to standard output
  • C. Write metrics to a custom SQL database
  • D. Send metrics to Application Insights as custom events
Show answer

Answer: A

Workspaces act as MLflow tracking servers, and mlflow.autolog() records parameters, metrics and models for supported frameworks with one line. Printing values only writes logs, a custom database adds work and loses the studio comparison view, and Application Insights is for runtime telemetry.

Want more questions like this? Full AI-300 practice tests →

Question 6. A sweep job uses Bayesian sampling over three hyperparameters. You add a bandit early termination policy, and the job fails validation. Why?

  • A. Bandit policies require at least five hyperparameters
  • B. Bayesian sampling does not support early termination policies
  • C. Early termination only works on serverless compute
  • D. The primary metric must be logged with Application Insights
Show answer

Answer: B

Bayesian sampling chooses new values based on previous results and does not support early termination policies. Grid and random sampling do. The number of hyperparameters and the compute target are not the cause.

Want more questions like this? Full AI-300 practice tests →

Question 7. Several training runs of the same model have completed with different settings. You need to choose the best run by validation accuracy and compare the runs' metrics side by side. Where do you do this most directly?

  • A. In the metrics of the managed online endpoint
  • B. In a data drift monitor
  • C. In the job comparison view of the studio, or with MLflow search_runs
  • D. In the Responsible AI dashboard
Show answer

Answer: C

Runs tracked with MLflow can be compared in the studio’s job comparison view or queried with the MLflow client, sorted by a logged metric. Endpoint metrics describe deployed models, drift monitoring compares production data, and a Responsible AI dashboard analyses one model.

Want more questions like this? Full AI-300 practice tests →

Question 8. A model's predictions are used in a regulated decision. Before promotion, you must show which features drive its predictions and which subgroups it makes most errors on. What should you generate?

  • A. A data drift monitor on the training data
  • B. An MLflow metric comparison between runs
  • C. An AutoML leaderboard
  • D. A Responsible AI dashboard with interpretability and error analysis
Show answer

Answer: D

The Responsible AI dashboard combines interpretability, which shows feature importance, with error analysis and fairness assessment across cohorts. Data drift monitoring compares production data over time, MLflow metrics give aggregate scores, and an AutoML leaderboard ranks models on one metric.

Want more questions like this? Full AI-300 practice tests →

Question 9. A managed online endpoint deployment fails. The logs show the scoring script's init() function raising an error that it cannot find a module. Which fix is most likely?

  • A. Add the missing package to the deployment environment and redeploy
  • B. Increase the instance count of the deployment
  • C. Move 100% of traffic to the deployment
  • D. Choose a larger instance type
Show answer

Answer: A

A missing module during init() means the package is not installed in the deployment’s environment. Adding it to the environment’s conda file and redeploying fixes it. More instances, a traffic change or a larger SKU do not install missing packages.

Want more questions like this? Full AI-300 practice tests →

Question 10. Every night, 2 million documents in a storage folder must be scored by a registered model. Results are needed the next morning, and no one calls the model interactively. What is the most cost-effective deployment?

  • A. A managed online endpoint called from a script in a loop
  • B. A batch endpoint invoked on the input folder
  • C. A notebook on a compute instance that runs overnight
  • D. A Kubernetes online endpoint with autoscaling
Show answer

Answer: B

A batch endpoint scores large volumes of files asynchronously on compute that runs only for the job, which suits nightly scoring. A managed online endpoint runs continuously for low-latency requests. A compute instance and a looping notebook are neither scalable nor operationally sound.

Want more questions like this? Full AI-300 practice tests →

Design and implement a GenAIOps infrastructure

Question 11. A production chat deployment should keep the exact model version it was evaluated with until the team tests and approves a newer version. What should you configure on the deployment?

  • A. Automatic upgrade to the default version
  • B. A lower tokens-per-minute rate limit
  • C. A pinned model version with automatic upgrade disabled
  • D. A stricter content filter
Show answer

Answer: C

Pinning the model version and disabling automatic upgrades keeps the evaluated version until the team moves deliberately. Automatic upgrade to the new default changes behaviour without evaluation. Rate limits and content filters do not control the model version.

Want more questions like this? Full AI-300 practice tests →

Question 12. A regulated customer requires that prompts and responses are processed only within a specific geography. Which deployment choice should you avoid?

  • A. A global standard deployment
  • B. A regional standard deployment in an approved region
  • C. A data zone deployment within the required geography
  • D. A regional provisioned deployment in an approved region
Show answer

Answer: A

Global deployment types may process requests in any region worldwide, which conflicts with a geography requirement. Regional and data zone deployments keep processing within a region or a defined zone. The model family itself does not determine where processing happens.

Want more questions like this? Full AI-300 practice tests →

Question 13. After public network access is disabled on a Foundry resource, apps in a virtual network can no longer reach it. A private endpoint exists, but name resolution still returns the public address. What is missing?

  • A. Re-enable public network access
  • B. Create a larger model deployment
  • C. Distribute the API key to the apps
  • D. Link the private DNS zone for the Foundry resource to the virtual network
Show answer

Answer: D

A private endpoint only works when DNS resolves the service name to its private IP, which requires linking the private DNS zone to the virtual network. Re-enabling public access or adding API keys defeats the purpose, and a larger deployment does not affect networking.

Want more questions like this? Full AI-300 practice tests →

Question 14. Two system prompt variants are proposed for a summarisation feature. How should the team decide which one to ship?

  • A. Chat with each variant a few times and pick the one that feels better
  • B. Run both against the same test dataset with the same evaluators and compare scores
  • C. Choose the shorter prompt to save tokens
  • D. Deploy both and let users pick one at random, without measuring
Show answer

Answer: B

Running both variants against the same versioned test set and scoring them with the same evaluators gives a fair, repeatable comparison. Manual chats are anecdotal, prompt length says nothing about quality, and shipping both at random without measurement produces no decision.

Want more questions like this? Full AI-300 practice tests →

Implement generative AI quality assurance and observability

Question 15. A support bot's answers are well written and address the question, but the legal team finds that some contain promises not found in any policy document. Which evaluator measures this problem?

  • A. Fluency
  • B. Relevance
  • C. Groundedness
  • D. Coherence
Show answer

Answer: C

Groundedness checks whether the response is supported by the provided context, so claims absent from the policy documents lower the score. Fluency, relevance and coherence can all be high for invented content.

Want more questions like this? Full AI-300 practice tests →

Question 16. Before launch, the security team wants evidence that a public-facing assistant resists attempts to make it produce violent or self-harm content. What should you run?

  • A. Risk and safety evaluations over an adversarial test dataset
  • B. Groundedness and relevance evaluations over normal questions
  • C. A load test measuring latency
  • D. A token consumption report
Show answer

Answer: A

Risk and safety evaluators score responses for harmful content categories, and running them over adversarial test inputs shows how the system behaves under attack. Quality evaluators do not measure harm, latency tests measure speed, and token reports measure cost.

Want more questions like this? Full AI-300 practice tests →

Question 17. Monthly model costs for an app doubled without an increase in users. You need to find which feature or step consumes the extra tokens. What gives the most direct answer?

  • A. The average latency chart for the deployment
  • B. The count of failed requests
  • C. Purchasing more provisioned throughput units
  • D. Traces with token usage per model call, grouped by feature or step
Show answer

Answer: D

Tracing records token counts on each model call span, and correlating them by feature or step shows where consumption grew. Latency and error metrics do not show token use. A bigger PTU commitment changes billing, not consumption. A quality evaluation measures answers, not cost.

Want more questions like this? Full AI-300 practice tests →

Optimize generative AI systems and model performance

Question 18. Users of an internal RAG app searching for error codes such as E-2231 get unrelated results, while descriptive questions work well. Retrieval uses vector search only. What should you implement?

  • A. A larger chat model
  • B. Hybrid search combining keyword and vector retrieval
  • C. A higher model temperature
  • D. Smaller chunks with no overlap
Show answer

Answer: B

Hybrid search adds keyword matching, which finds exact tokens such as error codes, and merges it with vector results that capture meaning. A bigger chat model or temperature change does not affect retrieval, and smaller chunks do not make an embedding match an exact code.

Want more questions like this? Full AI-300 practice tests →

Question 19. A RAG app frequently replies that it cannot find an answer, even when the relevant passage exists in the index and appears in the top results when you inspect them manually. What should you adjust first?

  • A. Fine-tune the chat model on example answers
  • B. Replace the embedding model and re-embed the corpus
  • C. Lower the similarity threshold for retrieved chunks
  • D. Remove the system prompt
Show answer

Answer: C

If relevant passages are retrieved but then dropped, the similarity threshold is filtering them out, so lowering it lets valid matches through. Fine-tuning and a new embedding model are larger changes aimed at other problems, and removing the system prompt would weaken grounding.

Want more questions like this? Full AI-300 practice tests →

Question 20. A team wants a model to reliably prefer concise, polite answers over technically correct but curt ones. It has many prompts, each with a preferred and a rejected response. Which fine-tuning method matches this data?

  • A. Direct preference optimisation (DPO)
  • B. Supervised fine-tuning on the rejected responses
  • C. Reinforcement fine-tuning without a grader
  • D. Fine-tuning the embedding model
Show answer

Answer: A

Direct preference optimisation trains on pairs of preferred and rejected responses to steer the model towards the preferred style. Supervised fine-tuning uses single ideal responses, reinforcement fine-tuning needs a grader that scores outputs, and changing the embedding model affects retrieval only.

Want more questions like this? Full AI-300 practice tests →

How did you do?

Sixteen or more correct suggests you are close. Below fourteen, the domain guides in this section are the fastest way to find the gaps. Questions 11–20 cover the generative AI half of the exam: if you missed more than three of them, spend your next study week on GenAIOps infrastructure, quality and observability and optimisation, because that is where AI-300 differs most from DP-100.