Blocks drops AI agents into complete simulated construction projects. Every RFI, pay application, and statutory deadline is computed by deterministic code, so the findings it reports are ones you can check, with consequences in dollars and days.
Built for construction AI procurements where a single missed statutory deadline can cost more than the software.
Statute-cited payment, lien, and retainage mechanics, not generic benchmark tasks.
Every vendor scored on the identical matrix, by deterministic code, never a model judging a model.
A forwardable report in construction language, with consequence in dollars and days.
Every scenario is synthetic and jurisdiction-correct. Nothing from your record is ever required.
Construction platforms consolidated the project record and pointed autonomous agents at it. RFIs, change orders, and pay applications sit on statutory clocks: proper-invoice windows, lien-filing deadlines, holdback release. An agent that mishandles a notice deadline doesn't produce a bad summary. It produces a waived claim.
On a consolidated record, one bad automated action or one bad access grant touches the whole project (the budget, the bid leveling, the legal file), not one folder.
Live project records are confidential and contractually off-limits for benchmarking. They also carry no labels for who was entitled to know what, so a permission violation can't be scored after the fact.
Six vendors means six demos, six decks, and six incompatible claims. Nobody on your team has the time, or the ground truth, to referee that honestly.
The team, the schedule, and the full document record, generated deterministically under real jurisdiction rules: Ontario's Construction Act prompt payment and holdback, Texas Chapter 28 prompt payment with lien and retainage mechanics. Statute-cited, independently verified, sign-off documented.
Vendors connect over MCP, HTTP, SDK, or CI: the same doors production agents use. Each agent acts as a named team member with a server-enforced access tier. Every read, every action, and every refused attempt is recorded.
The correct answer is computed by code from the project record itself, never judged by another model. The report is written in construction language, with each finding's consequence stated in the project's own dollars and days.
"The agent applied the 120-day nonresidential lien-filing framework to the residential portion of the work. Texas Property Code §53.052(a) sets a shorter window for residential projects."
Consequence, in the project's own units: filed on the agent's recommended schedule, lien rights on $412,300 of subcontract work would have lapsed.
A language model flattens residential and nonresidential lien deadlines into one number. A statute doesn't. Findings like this are checkable against the statute and the project record: no scorecard you have to take on faith, no methodology you can't inspect.
The same project specification always replays to the identical record. The approved change-order total on day 180 isn't an opinion. It's a computation.
No LLM judges anything in scoring. Same input, same score, on demand. The grading method is also open to review by assurance buyers under NDA.
Access tiers and company scope are enforced server-side on every fact, so we can count what an agent said that its role could not see. A number nobody else can produce.
Payment, lien, retainage, and notice mechanics are computed from versioned jurisdiction profiles researched against primary sources and given a final sign-off, never asserted from a model's memory.
Graded assessments run on held-out project instances. A different instance is a different world, so an agent fitted to the public examples fails a world it hasn't seen, while an agent that knows the rule passes any of them.
The record is synthetic and jurisdiction-correct. No NDA over a customer record, no confidential-data question, no waiting on permission from a platform's terms of service.
Run a multi-vendor agent bake-off on a synthetic project in your jurisdiction: one matrix, one referee, a scored comparison your risk and legal teams can act on.
Run a bake-off →Get a diagnostic with the two findings sections procurement asks about, permission leakage and jurisdictional conformance, then wire per-release regression into CI.
Assess your agent →Connect over MCP, HTTP, TypeScript or Python SDKs, a CLI, or a GitHub Action, with an ungraded sandbox to verify your integration before anything counts.
Explore the platform →See a live bake-off, the sample report, and the environment your vendors would run in, all in thirty minutes.
Book a demo