Free AI-300 sample questions with answers
Here are five free AI-300 sample questions, one from each skill area, with the answer and the reasoning behind it. Try each one before opening the answer. These are original scenario questions written against the official skills outline, not questions from the real exam.
Question 1. A GitHub Actions workflow deploys Azure Machine Learning resources with Bicep. The security team forbids storing any long-lived Azure credential in GitHub. How should the workflow authenticate to Azure?
- A. Store a service principal client secret as a GitHub secret
- B. Use OpenID Connect with a federated identity credential on a Microsoft Entra app or managed identity
- C. Store the storage account key of the workspace as a GitHub secret
- D. Use a GitHub personal access token in the azure/login step
Show answer
Answer: B
OpenID Connect with a federated identity credential lets the workflow exchange a short-lived GitHub token for an Azure token, so no secret is stored. A client secret or an account key is a long-lived credential stored in GitHub. A personal access token authenticates to GitHub, not to Azure.
Want more questions like this? Full AI-300 practice tests →
Question 2. A new model version is deployed to an existing managed online endpoint as a second deployment. You want 10% of live traffic to reach it, and to be able to return all traffic to the current version immediately if error rates rise. What should you do?
- A. Create a new endpoint for the new version and update all clients
- B. Update the existing deployment in place with the new model
- C. Set the endpoint traffic to 90% for the current deployment and 10% for the new one
- D. Deploy the new version to a batch endpoint and compare outputs
Show answer
Answer: C
An endpoint can host several deployments and split traffic between them by percentage. Setting 90/10 exposes the new version to a slice of users, and setting it back to 100/0 is an instant rollback without redeploying. A new endpoint changes the scoring URI for clients, replacing the model in place removes the fallback, and a batch endpoint does not serve live requests.
Want more questions like this? Full AI-300 practice tests →
Question 3. A foundation model deployed in Microsoft Foundry serves a customer-facing chat app at a steady, high request volume. Latency has become unpredictable at peak times, and finance wants a predictable monthly cost. What should you configure?
- A. A provisioned throughput deployment sized in PTUs
- B. Switch to the smallest model in the same family
- C. Increase the client retry count on throttled requests
- D. Cache model responses in the user’s browser
Show answer
Answer: A
Provisioned throughput units reserve model processing capacity, which gives consistent latency at high, steady volume and a predictable cost. A smaller model changes quality rather than capacity. Retries add load at exactly the wrong moment. Caching in the browser does not help unique chat requests.
Want more questions like this? Full AI-300 practice tests →
Question 4. A RAG assistant sometimes states facts that do not appear in the retrieved documents, even though its answers are fluent and on topic. Which evaluation metric targets this problem most directly?
- A. Fluency
- B. Relevance
- C. Coherence
- D. Groundedness
Show answer
Answer: D
Groundedness measures whether the response is supported by the provided context, so unsupported claims lower it. Fluency measures language quality and relevance measures whether the answer addresses the question; both can be high while the answer is invented. Coherence measures how well the answer hangs together.
Want more questions like this? Full AI-300 practice tests →
Question 5. A support search in a RAG app misses documents when users type exact product codes such as XR-4410, although it handles natural-language questions well. Retrieval uses vector search only. What should you change first?
- A. Deploy a larger chat model
- B. Switch retrieval to hybrid search combining keyword and vector queries
- C. Increase the chunk size to include more text per chunk
- D. Lower the similarity threshold so more chunks are returned
Show answer
Answer: B
Hybrid search runs keyword and vector queries together and merges the results, so exact tokens such as product codes are matched by the keyword side while meaning is matched by the vector side. A larger model does not change what is retrieved. Bigger chunks dilute the embeddings further, and lowering the similarity threshold returns more loosely related results without matching the code.
Want more questions like this? Full AI-300 practice tests →
How did you do?
Questions 1 and 2 are the operational half of the exam and feel familiar to anyone who ran DP-100 labs in a pipeline. Questions 3 to 5 are the generative AI half. If those slowed you down while the rest did not, you are where most former DP-100 candidates are. Week 3 of the study plan covers Foundry, and the quality and observability guide covers evaluators.
For a longer check, take the 20-question practice test.