Subodh Jena
BlogNotes

Blog

Writing

Life, tech, and everything in-between.

I Read the Code

Jul 3, 2026·2 min·Comic

The planner, coder, judge split is a good harness: an expensive model plans, a cheap one writes, the expensive one judges, with the budget spent where judgment lives. But an LLM judge is a filter, not an alibi; it catches the cheap model's misses, not whether the plan solved your problem or whether the edge cases the models quietly agreed on match the ones your users will find. The one step that does not delegate is the last: somebody opens the diff and reads it.

Agent Evaluation: LLM-as-Judge, Pass-at-K, and Benchmarks

May 6, 2026·8 min·AI

An agent that cannot be measured cannot be improved. Evaluation starts with twenty queries and an LLM-as-judge, scales up through pass-at-k metrics and standard benchmarks, and never trusts any single layer alone.

Writing

BlogNotes

Work

ExperimentsPortfolio

Connect

AboutContact

© 2026 Subodh Jena

X (Twitter)GitHubLinkedIn