Skip to content
DevMeme
7492 of 7592
AI ML Post #8212 · source on Telegram

The Evaluation Metric Is Reading the Code

Description

A dark-mode X thread begins with verified developer Mitchell Hashimoto (@mitchellh, "7h") writing, "I'm having a lot success using Fable xhigh as a planner/architect, using GPT 5.5 xhigh (subscription) as a coder, then Fable xhigh again as a judge. At API pricing, planning+judge costs are in the ~few dollar range compared to typical $50+ full round trips." The post continues, visibly truncated by "Show more," with "I've seen some others using dumber/cheaper coders, but GPT 5.5 even at xhigh," and its UI shows 195 replies, 152 reposts, 3.4K likes, and 176K views. A reply from "~~ Pooja ~~" (@poojabnf, "6h") asks, "What metrics are you using to measure the success of this approach, and how do you handle potential errors or inconsistencies introduced by the AI tools?" and shows 2 replies and 31 likes. Hashimoto's deadpan answer in the embedded X.com post is simply "I read the code," timestamped "10:43 · 02/07/2026 · 142K Views," with 22 replies, 129 reposts, 1.3K likes, and 75 bookmarks; the exchange grounds an elaborate multi-model planner/coder/judge workflow in ordinary human code review and engineering judgment.

Comments

1
Anonymous ★ Top Pick The final evaluator is a biological model with repository access and a coffee-based inference budget.
  1. Anonymous ★ Top Pick

    The final evaluator is a biological model with repository access and a coffee-based inference budget.

Use J and K for navigation