Collate Automated and Manual Scores
Start with two complete sets of scores: one from your automated LLM-as-judge and one from your manual 'vibe check'. Ensure you have parallel scores for each model's output on every task.
Start with two complete sets of scores: one from your automated LLM-as-judge and one from your manual 'vibe check'. Ensure you have parallel scores for each model's output on every task. My actual values: [insert the files, settings, accounts, or constraints for this step] Give me the exact commands, settings, or output to use. Finish with a pass or fail check for collate automated and manual scores.




