AI NewsAI4Bharat Builds 4,000-Test Benchmark, Finds AI Judges Fail Half the Time
AI4Bharat Builds 4,000-Test Benchmark, Finds AI Judges Fail Half the Time
12:51 PM IST · July 23, 2026

AI4Bharat has created FOCUS, a benchmark designed to measure how well evaluator VLMs detect mistakes across both image-to-text (I2T) and text-to-image (T2I) tasks.
read more