1. Select concrete tasks. Define verifiable outputs: extract a field, classify a request or summarize a decision. Write acceptance criteria before running calls. Separate simple tasks from ambiguous cases; do not rely only on a score given by the model itself.
2. Prepare synthetic examples. Reproduce document formats, lengths and difficulty without copying real identities. Include missing fields and hostile instructions in text. Reserve a separate set to check performance on examples not used during tuning.
3. Compare identical inputs and limits. Record the model, available version, policy, context size, latency and consumption. Distinguish incorrect answers, incomplete answers and appropriate refusals. Longer output is not evidence of better performance.
4. Evaluate privacy alongside quality. Compare usefulness with placeholders, removal and blocking. Check whether responses reconstruct identities or repeat unnecessary details. Choose the combination meeting your criteria and repeat evaluation when models, prompts or formats change.