AI-300 GenAI quality and observability
Implement generative AI quality assurance and observability is worth 10–15% of AI-300. It asks two questions about a generative AI application or agent: is it good enough to ship, and can you see what it is doing once it runs? The first is answered with evaluations, the second with monitoring, tracing and cost tracking in Microsoft Foundry.
Evaluation
Test datasets and data mapping
An evaluation runs over a dataset, typically a JSONL file where each row holds a query, the response, the retrieved context and sometimes a ground-truth answer. Evaluators expect named inputs, so data mapping tells each evaluator which column is the query, which is the response and which is the context. A wrong mapping is a classic reason for evaluations that run but produce meaningless scores.
A useful test set covers normal questions, edge cases, questions the system should decline, and adversarial inputs. It is versioned alongside the prompts it tests.
Quality metrics
| Metric | Question it answers | Needs |
|---|---|---|
| Groundedness | Is the answer supported by the provided context? | Response and context |
| Relevance | Does the answer address the question? | Query and response |
| Coherence | Does the answer read logically and hang together? | Query and response |
| Fluency | Is the language well formed? | Response |
| Similarity and F1-style metrics | How close is it to a known correct answer? | Ground truth |
Groundedness is the metric most often tested, because it is the one that catches hallucination in RAG. A fluent, relevant answer can still be ungrounded.
Most quality evaluators are AI-assisted: a judge model scores each row. That makes them flexible but means the judge deployment and its cost are part of the setup.
Risk and safety evaluations
Risk and safety evaluators score content for categories such as hateful and unfair content, sexual content, violence and self-harm, and check for issues such as protected material or susceptibility to indirect prompt injection. They are typically run against both normal and adversarial test data, the latter often generated with a simulator, before release.
Automated evaluation workflows
Built-in evaluators cover the common metrics; custom evaluators, written as code or as a prompt for a judge model, cover domain rules such as “always cite a policy number”. The operational step is to run evaluations automatically: in a GitHub Actions workflow on every prompt or model change, with thresholds that fail the build, and continuously on sampled production traffic.
Observability
Continuous monitoring
Foundry can evaluate a sample of production traffic on a schedule, so quality and safety scores are tracked over time instead of only at release. A drop in groundedness after a data source changes shows up there first.
Performance and cost
Watch latency, throughput and response time per deployment, plus throttled requests. On cost, the unit is the token: prompt tokens and completion tokens per request, per user and per feature. Common optimisations include trimming the retrieved context, shortening system prompts, capping output length, caching repeated answers and routing simple requests to a smaller model.
Tracing and logging
Tracing records each step of a request as spans: the incoming query, retrieval, each model call with its tokens, and each tool call an agent makes. Foundry’s tracing is built on OpenTelemetry and sends traces to Application Insights, where they can be queried and correlated with the rest of the application. Traces are the tool for answering “why did this particular answer go wrong?”; aggregate metrics are not.
Sample questions
Question 1. You run a groundedness evaluation on a JSONL dataset with columns question, answer and sources. The evaluation completes, but every score is at the minimum, even for answers you know are correct. What is the most likely cause?
- A. The judge model deployment is too small
- B. Groundedness requires a fluency evaluator to run first
- C. The evaluator’s context input is not mapped to the sources column
- D. The dataset has too few rows
Show answer
Answer: C
Evaluators read named inputs, so if the context input is not mapped to the sources column, the evaluator sees no context and judges every answer as unsupported. The judge model’s size does not explain uniformly minimum scores. Fluency is a different metric, and more rows would not change the per-row result.
Want more questions like this? Full AI-300 practice tests →
Question 2. Every change to a RAG application's prompts must be blocked from reaching production if answer quality drops. Which approach fits best?
- A. Ask a reviewer to try a few questions after each change
- B. Add an automated evaluation step with metric thresholds to the deployment workflow
- C. Rely on continuous monitoring in production to catch regressions
- D. Configure an alert when average latency increases
Show answer
Answer: B
An evaluation step in the deployment workflow scores each change on the same test set and fails the run if metrics such as groundedness fall below a threshold, blocking the release automatically. Manual spot checks are not consistent, production monitoring detects problems only after release, and latency alerts do not measure answer quality.
Want more questions like this? Full AI-300 practice tests →
Question 3. A user reports that an agent gave a wrong answer yesterday afternoon. Aggregate metrics look normal. You need to see which documents were retrieved and which tool calls the agent made for that request. What should you use?
- A. The token consumption report for the deployment
- B. The average latency chart in Foundry monitoring
- C. A new evaluation run over the test dataset
- D. The trace for that request in Application Insights or Foundry tracing
Show answer
Answer: D
Traces record each step of an individual request as spans, including retrieval results and tool calls, so the failing request can be inspected. Aggregate dashboards and token cost reports summarise many requests. A new evaluation run tests current behaviour, not what happened yesterday.
Want more questions like this? Full AI-300 practice tests →
What to practise
Build a 20-row test set for a small RAG app, map its columns, and run groundedness, relevance, coherence, fluency and one safety evaluator. Break the retrieval on purpose and watch groundedness fall. Then enable tracing, send a few requests, and find one request’s retrieval and model spans in Application Insights.