AI-300 GenAI optimization: RAG and tuning
Optimize generative AI systems and model performance is worth 10–15% of AI-300. It covers the two main ways to make a generative AI application better once it works: improve retrieval in a RAG system, and customise the model itself through fine-tuning. The exam asks which lever fits a symptom, so the skill is diagnosis first, then the right fix.
Optimising RAG
Most RAG problems are retrieval problems. If the right passage never reaches the model, no prompt will fix the answer. The outline names four areas.
Chunking, thresholds and retrieval strategy
| Setting | Too small | Too large |
|---|---|---|
| Chunk size | Passages lose context; answers are fragmentary | Embeddings blur several topics; irrelevant text fills the context |
| Chunk overlap | Sentences split across boundaries are lost | Duplicate content wastes tokens |
| Top-k | The relevant passage is missed | Noise and cost rise; the model may cite the wrong passage |
| Similarity threshold | Weak matches pass through | Valid matches are filtered out and the app answers “I don’t know” |
Chunk along document structure — headings, sections, paragraphs — where possible, rather than a fixed character count. Other retrieval strategies include query rewriting, metadata filters, and re-ranking the top results with a stronger model.
Embedding models
The embedding model decides what “similar” means. A general model can struggle with specialised vocabulary in legal, medical or internal product language. Options, in order of effort: try a different or larger embedding model, compare them on a labelled retrieval set, and fine-tune an embedding model on domain pairs where the gain justifies it. Changing the embedding model means re-embedding the whole corpus; the index and the query must use the same model and dimensions.
Hybrid search
Vector search finds meaning; keyword search finds exact tokens such as product codes, error numbers and names. Hybrid search runs both and merges the ranked lists, in Azure AI Search with Reciprocal Rank Fusion. Adding the semantic ranker on top re-scores the merged results with a language model. Hybrid is the usual answer when users search with exact identifiers that pure vector search misses.
Measuring and A/B testing
Retrieval quality is measured with relevance metrics over a labelled set: did the right document appear, and how high? Answer-level evaluators such as groundedness and relevance measure the end result. Changes are compared with A/B tests, sending a share of traffic or a fixed test set to each variant and comparing scores, not by trying a few questions by hand.
Fine-tuning and customisation
When to fine-tune
Fine-tuning changes how a model behaves: tone, format, adherence to a task, or skill on a narrow domain. It does not reliably teach new facts that change often; that is what RAG is for. The usual order is prompt engineering, then RAG, then fine-tuning when the first two plateau.
Methods
| Method | Training data | Use it for |
|---|---|---|
| Supervised fine-tuning (SFT) | Example prompts with ideal responses | Teaching a format, style or task |
| Preference optimisation, such as DPO | Prompts with a preferred and a rejected response | Steering towards answers people prefer |
| Reinforcement fine-tuning | Prompts plus a grader that scores responses | Improving reasoning on tasks with checkable answers |
Parameter-efficient approaches such as LoRA train a small set of additional weights instead of the whole model, which cuts cost and time.
Synthetic data
When real examples are scarce, a stronger model can generate synthetic training data. It must be reviewed and filtered: duplicates, errors and unsafe content are learned just as eagerly as good examples. Keep a real, held-out test set that no synthetic data touches, or you measure the model on its own homework.
From development to production
A fine-tuned model is trained, evaluated against the base model on the same test set, deployed, and monitored like any other deployment. Watch training and validation loss for overfitting during training, and quality, safety and cost after deployment. Version the training data and the resulting model together, so any production model can be traced back to what it was trained on.
Sample questions
Question 1. A RAG assistant over long technical manuals often returns answers that mix instructions from unrelated sections. Retrieved chunks are 3,000 tokens each. What should you try first?
- A. Reduce chunk size and split along section headings
- B. Deploy a larger chat model
- C. Increase top-k to retrieve more chunks
- D. Remove chunk overlap entirely
Show answer
Answer: A
Very large chunks produce embeddings that blur several topics, so unrelated content is retrieved together. Smaller chunks aligned to section boundaries make each chunk about one thing. A larger chat model or higher top-k adds more of the same mixed context, and removing overlap does not address topic mixing.
Want more questions like this? Full AI-300 practice tests →
Question 2. A customer service model answers correctly but ignores the company's required response structure and tone, even with detailed instructions and examples in the prompt. The team has 2,000 approved example conversations. What should you do?
- A. Add the style guide to the RAG index
- B. Switch to a larger embedding model
- C. Fine-tune the model with supervised fine-tuning on the approved conversations
- D. Increase the temperature of the model
Show answer
Answer: C
Supervised fine-tuning on approved examples teaches a consistent format and tone once prompting has plateaued. Adding more documents to the index affects facts, not style. A bigger embedding model changes retrieval. Raising the temperature makes outputs more varied, not more consistent.
Want more questions like this? Full AI-300 practice tests →
Question 3. You generate 10,000 synthetic training examples with a large model to fine-tune a smaller one. The fine-tuned model scores very highly on a test set that was generated in the same run, but performs poorly with real users. What is the most likely cause?
- A. The fine-tuning job used too few epochs
- B. The test set was not independent real data, so it overstated performance
- C. The smaller model has too small a context window
- D. The model should have been deployed with provisioned throughput
Show answer
Answer: B
Evaluating on synthetic data from the same generation run measures how well the model imitates that data, not real-world performance. A held-out test set of real examples is needed. The number of epochs, model size and the deployment type do not explain the gap between the two test results as directly.
Want more questions like this? Full AI-300 practice tests →
What to practise
Index the same documents twice with different chunk sizes, and compare vector-only, keyword-only and hybrid search on ten questions you know the answers to. Then prepare a small supervised fine-tuning file, run a fine-tuning job, and compare the tuned model with the base model on a held-out set of real examples.