AI-300 GenAIOps infrastructure explained
Design and implement a GenAIOps infrastructure is worth 20–25% of AI-300. It is the generative AI counterpart of the MLOps infrastructure domain: set up Microsoft Foundry securely and from code, deploy foundation models in the way that fits the workload, and treat prompts as versioned source rather than text pasted into a portal.
Foundry resources and projects
A Foundry resource is the Azure resource that holds model deployments, connections and security settings. Projects live inside it and give teams a workspace for agents, evaluations and files, sharing the resource’s deployments and connections. A typical pattern is one resource per environment or business unit, with projects per application or team.
Like Azure Machine Learning workspaces, Foundry resources and projects can be declared in Bicep and deployed with the Azure CLI, which is how dev, test and prod stay identical.
Identity and access
Two things need identities:
- People get Azure RBAC roles on the resource or project. Separate roles let someone build and test without being able to change infrastructure or deployments.
- Applications and the resource itself use managed identities. An app calling a model uses its managed identity and Microsoft Entra authentication instead of an API key; the Foundry resource uses its own identity to reach connected services such as storage or Azure AI Search.
Keys still exist, but a scenario that mentions rotation, auditing or “no secrets in configuration” is pointing at managed identities and RBAC.
Private networking
Private endpoints give private access to the Foundry resource from a virtual network, and public network access can be disabled. As with Machine Learning, connected resources such as storage and search need private access as well, and DNS must resolve the private endpoints. The failure pattern is familiar: everything works from the portal on a public network, and nothing works after public access is turned off.
Deploying foundation models
Deployment options
| Option | What it means |
|---|---|
| Serverless API | Pay per token; Microsoft hosts the model; no infrastructure to manage |
| Managed compute | The model runs on dedicated VMs in your subscription; you pay for the compute |
| Standard deployment | Pay per token on shared capacity, subject to rate limits |
| Provisioned throughput | Reserved capacity bought in provisioned throughput units (PTUs) |
Standard and provisioned deployments also come in regional, data zone and global variants, which control where requests may be processed. A data residency requirement rules out global deployments.
When PTUs pay off
Pay-per-token is cheapest for spiky or low volume. Provisioned throughput suits high, steady volume where predictable latency and a predictable bill matter more than paying only for what you use. Undersized provisioned deployments return throttling responses when capacity is exhausted, so sizing and spillover planning are part of the objective.
Choosing and versioning models
Model selection weighs task fit, quality on your own data, context window, latency, cost and region availability. The model catalog and benchmarks narrow the field; your own evaluation decides. Deployments pin a model version and have an upgrade policy: upgrade automatically when a new default version appears, or stay pinned until you test and move. Production deployments should change version deliberately, after evaluation.
Prompts as source code
The outline has three prompt objectives: design and develop prompts, create variants and compare them, and version them in Git.
- Design: a clear system message, examples where they help, explicit output format, and grounding instructions for RAG.
- Variants: two or more versions of the same prompt run against the same test data and scored with the same evaluators. The winner is chosen on numbers, not on a few manual chats.
- Versioning: prompts stored as files, for example in a templated prompt format, change through pull requests and are deployed by the same workflow as the code. A regression can then be traced to one commit and reverted.
Sample questions
Question 1. A web app calls a model deployed in Microsoft Foundry using an API key stored in its app settings. Security requires removing the key and auditing access per application. What should you implement?
- A. Rotate the API key every 30 days
- B. Enable a managed identity for the app and grant it a role on the Foundry resource
- C. Move the API key to Azure Key Vault and reference it from app settings
- D. Restrict the Foundry resource to the app’s outbound IP addresses
Show answer
Answer: B
Giving the app a managed identity and a role on the Foundry resource lets it authenticate with Microsoft Entra ID, removing the key, and every call is attributed to that identity. Rotating the key or moving it to Key Vault keeps a shared secret. Restricting by IP address does not identify the application.
Want more questions like this? Full AI-300 practice tests →
Question 2. An internal assistant sends a steady, high volume of requests to a chat model throughout the working day. Users complain about variable latency at peak times, and finance wants a predictable monthly bill. Which deployment type fits best?
- A. Standard deployment
- B. Global standard deployment
- C. Global batch deployment
- D. Provisioned throughput deployment
Show answer
Answer: D
A provisioned throughput deployment reserves capacity in PTUs, giving consistent latency and a predictable cost for steady, high volume. Standard and global standard pay per token on shared capacity, so latency varies with demand. Batch deployments process asynchronous jobs, not interactive chat.
Want more questions like this? Full AI-300 practice tests →
Question 3. After a colleague edited a system prompt in the portal, answer quality dropped, and nobody can tell what changed or restore the previous wording. What should the team implement?
- A. Store prompts as files in Git, change them via pull requests, and evaluate variants before deploying
- B. Take screenshots of the prompt after every change
- C. Document the current prompt on the team wiki
- D. Pin the model version on the deployment
Show answer
Answer: A
Storing prompts as files in a Git repository, changed through pull requests and evaluated before deployment, gives history, review and a one-step revert. Screenshots and wiki pages are not enforced or deployable. Pinning the model version addresses model changes, not prompt changes.
Want more questions like this? Full AI-300 practice tests →
What to practise
Deploy a Foundry resource and project with Bicep, grant an app’s managed identity access, and call a deployed model without a key. Then store the system prompt in your repository, create a second variant, run both against the same small test set, and merge the better one through a pull request. That covers the domain end to end.