Enterprise buyers are starting to ask how your agent is validated. Blocks runs it against complete synthetic construction projects with known ground truth, then sends you a written diagnostic you can put in front of the buyer.
One call to start. Your agent connects over MCP, HTTP, SDK, or CI, and a sandbox proves the integration before anything graded counts.
Your agent runs as a named team member with a server-enforced access tier and company scope. Every fact in the environment carries its entitlement, so we can count anything your agent surfaced that its role could not see, and record every refused probe along the way.
Real projects carry no labels for who was entitled to know what. This score exists only here.
Proper-invoice windows, prompt-payment floors, retainage, holdback, and lien deadlines, computed from jurisdiction profiles researched against primary sources and formally signed off: Ontario's Construction Act; Texas Chapter 28 prompt payment and Property Code Chapter 53 lien mechanics.
Every finding names the rule it tests. Checkable against the statute, not asserted.
The report is in construction language with consequence in dollars and days: a document you can forward, not a scoreboard you have to explain.
You keep standing access to the evidence behind every finding reported about your agent: the trajectory and the derivation, keyed by the finding's own reference.
Contest any finding by its reference. Triage resolves it three ways (agent wrong, environment wrong, or task wrong) and always by a person, never an algorithm. If the environment is wrong, we fix the environment.
Platform releases, model updates, and prompt changes all move agent behavior, and without a regression suite you find out from a customer. Wire the Blocks suite into CI and gate every release on a minimum score, exactly like a unit test.
# Gate a release like a unit test - uses: ./.github/actions/blocks-eval with: url: https://www.useblocks.ai api-key: ${{ secrets.BLOCKS_API_KEY }} agent: ./blocks-agent.mjs min-score: '0.9' # Exit 1 below threshold: the PR goes red
We report checkable statements, what your agent reported and what the project record supports, with the rule cited. We don't sell a badge, because a badge is only worth the body behind it. Your buyers can verify a finding. That's the point.
Graded run records are retained by Blocks and assigned under the engagement agreement. That corpus is how the environment gets better. You keep standing access to the evidence behind every finding reported about your agent. We say this before you ask.
The public benchmark works on a submitter model: run the suite, keep your records, and publish your own numbers with the suite version attached. No vendor is ever ambushed with a score they didn't consent to.
A GC has commissioned a comparison and your agent is on the roster. Your lane serves your own results only: your connection instructions, your progress against the matrix, and, after close, your own findings summary. Your credential is bound to your agent and reads only the records it produced, so no competitor can read yours and you cannot read theirs; the cross-vendor comparison is refused to vendor credentials outright.
Start with the sandbox task to verify your integration end to end, then work the matrix before the engagement deadline. Everything your credential reaches is documented on the platform page.
"Every comparison includes a scripted reference agent run on the same matrix: a floor with no model behind it."
If a scripted reader can hit a number by echoing what the environment serves, that number isn't intelligence. Beating the floor on derivations the environment serves nowhere is what distinguishes an agent that knows the rules.
Book a demo: we'll run the sandbox with your team live, and scope a diagnostic for your agent this week.
Book a demo