Choose LLM Adaptation by the Failure You Need to Fix
September 29, 2026 · 7 min read

Teams often choose between prompting, retrieval-augmented generation (RAG), and fine-tuning as if they were competing implementation styles. They are not. Each changes a different part of the system.
Prompting changes the instructions and context supplied at inference time. RAG changes what evidence the model can access. Fine-tuning changes the model’s learned behavior. The practical question is therefore not “Which technique is best?” but “What is causing the current failures?”
That distinction matters because the wrong intervention can produce an impressive demo while leaving the underlying defect untouched. Fine-tuning will not keep policy documents current. A vector database will not reliably enforce an output schema. A longer prompt will not teach a small model a complex, recurring transformation it consistently fails to perform.
Diagnose errors before choosing an approach
Start with a representative evaluation set, not a technology decision. Include routine requests, edge cases, ambiguous inputs, adversarial phrasing, and examples where the correct answer is to abstain.
For every failed response, assign a primary cause:
- Instruction failure: The model had the necessary information but misunderstood the task, constraints, or output format.
- Knowledge failure: The answer depended on facts that were absent, private, or outdated.
- Behavior failure: The model repeatedly failed at a stable pattern despite clear instructions and sufficient context.
- Retrieval failure: Relevant evidence existed but was not found, ranked, filtered, or assembled correctly.
- Workflow failure: The task required tools, validation, permissions, or deterministic business logic outside the model.
This taxonomy prevents teams from treating every quality issue as a model issue. If pricing is wrong because an old catalog was retrieved, changing the model is wasted effort. If retrieved passages are correct but the model ignores them, retrieval alone is not the fix.
Measure results by failure category. A single aggregate accuracy score hides whether an intervention actually addressed its intended problem.
Use prompting for task clarity and fast control
Prompting is the right tool when the model can already perform the task but needs clearer direction. Typical cases include classification labels, response structure, tone, decision criteria, and rules for using supplied context.
Good prompts define the contract around the model. They state the objective, identify authoritative inputs, delimit untrusted content, specify the output schema, and explain when to abstain or escalate. Few-shot examples help when they clarify a boundary that prose does not express well.
Prompting is especially valuable early because it is cheap to change and easy to test. It exposes whether the task itself is well specified. If engineers cannot write a clear instruction and reviewers cannot agree on expected outputs, fine-tuning will encode disagreement rather than resolve it.
Prompting reaches its limit when instructions become a growing patchwork. Repeated exceptions consume context, interact unpredictably, and are difficult to maintain. Moving those rules into a fine-tuned model may help only if the desired behavior is stable and sufficiently represented in training data. Otherwise, simplify the task or implement deterministic validation.
Use RAG when answers depend on changing evidence
RAG is appropriate when the model must answer from information that is current, proprietary, traceable, or too large to include in every request. Examples include support documentation, product catalogs, contracts, incident histories, and internal engineering standards.
The defining advantage is not merely access to more text. It is the ability to update knowledge without retraining and to show the evidence behind an answer.
A production RAG system needs more than embeddings and a vector store. It needs:
- ingestion with ownership, timestamps, and access-control metadata;
- segmentation aligned with document structure and likely questions;
- hybrid retrieval where exact terms, identifiers, or error codes matter;
- ranking and filtering that preserve tenant and user permissions;
- citations linked to the specific claims they support;
- tests for answer quality and retrieval quality separately.
The last point is crucial. Evaluate whether the required evidence appeared in the retrieved set before judging generation. If it did not, tune ingestion, queries, filters, or ranking. If it did, inspect whether the prompt makes evidence authoritative and whether conflicting sources are resolved.
Do not use RAG to store behavioral rules that should always apply. A safety constraint retrieved only when semantically similar to a query is not a dependable constraint. Put mandatory controls in application logic, system instructions, and validation layers.
Use fine-tuning for repeated, stable behavior
Fine-tuning is justified when the desired capability is demonstrated by many high-quality examples, remains stable over time, and is expensive or unreliable to express in each prompt.
Strong candidates include domain-specific classification, consistent transformation between formats, specialized extraction, and a recurring response style with subtle conventions. Fine-tuning can also make a smaller model viable, reducing inference cost or latency at sufficient volume.
It is a poor substitute for a knowledge base. Facts embedded in weights are difficult to update, cite, delete, or scope by user permissions. Fine-tuning also introduces operational obligations: dataset versioning, privacy review, training pipelines, model registry controls, regression testing, and rollback procedures.
Before training, establish that the base model fails for behavioral reasons. Then compare the tuned model with a strong prompted baseline on a held-out set. Include rare cases and negative examples, not just the polished examples used to justify the project.
The economics matter. Training cost is often smaller than the cost of curating examples and maintaining evaluations. Fine-tuning makes sense when that investment is offset by measurable gains in quality, latency, context usage, or unit cost.
Combine techniques around system boundaries
Many production systems need more than one technique, but each component should have a clear job.
Consider an assistant that drafts responses to enterprise support tickets. RAG retrieves the current product documentation, customer entitlements, and known incidents. The prompt defines the task, requires citations, and instructs the model not to invent remediation steps. A fine-tuned model may classify ticket intent or transform the draft into the company’s established support format. Application code checks permissions, redacts secrets, and blocks unsupported actions.
This is better than asking one adaptation method to do everything. It also makes failures easier to locate.
A useful decision sequence is:
- Can the issue be fixed by making the task and output contract explicit?
- Does the response require evidence that changes, is private, or must be cited?
- Does a stable behavior still fail across a substantial, representative dataset?
- Should the requirement be deterministic code rather than probabilistic model behavior?
Run controlled experiments rather than stacking techniques immediately. Change one layer, rerun the same evaluation set, and compare improvements by failure category, latency, and cost.
Account for operational cost, not just answer quality
Architecture reviews should compare the full production burden.
Prompting adds context tokens, prompt versioning, and regression risk when instructions change. RAG adds ingestion jobs, search infrastructure, document governance, permission enforcement, and retrieval observability. Fine-tuning adds training data governance, model lifecycle management, and the possibility of behavior drifting between model versions.
Track at least task success, unsupported-claim rate, abstention quality, p95 latency, and cost per completed task. For RAG, add retrieval recall and citation correctness. For fine-tuning, compare gains against the prompted base model and monitor regressions outside the target behavior.
Choose the smallest operational surface that resolves the measured failure. Complexity should be purchased with evidence.
Takeaway
Prompting clarifies the job. RAG supplies current and governed evidence. Fine-tuning teaches stable, repeated behavior. Diagnose failures first, test one intervention at a time, and keep permissions, validation, and business rules outside the model where deterministic control is required.