Built on data, not opinions
- organisations analysed · - active contributors · - regions
Then: count the output. Now: measure the steering.
Legacy metrics count what the machine produced - lines, coverage, rule hits. AgenticReview measures what the humans did with it: guardrails set before the agent starts, plans written before it builds, review and correction after it delivers.
How the model was built
A reference cohort of 1,505 public engineering organisations was selected - 215 in each of seven regions - and is analysed continuously. The data was explored iteratively to find what actually separates software that can be owned and changed from software that cannot. Calibration anchors were set from that reference. 34 signals from git history become eight meters, scored on levels 1-5; the meters combine into four composite indices, and the indices into one total from 0 to 10.
Quality scored, stance mapped
Only what no competent architect would argue against is scored - dead code, duplicated rules, missing tests. Strategic choices where both sides have defenders, such as strict vs flexible schema, are mapped, never scored.
Versioned and recalibrated
Anchors are recalibrated as the benchmark grows. Every report states its model version, so results stay comparable over time.