In 2026, construction platforms consolidated the project record and pointed autonomous agents at it. That made agents genuinely useful, and it made every agent mistake project-wide, on documents that are contractually and statutorily binding. Nobody could prove an agent was safe to turn on. So we built the environment where that proof is possible.
The problem isn't that construction AI is bad. It's that the industry has no way to know. Live project records are confidential and contractually off-limits for benchmarking. Real projects carry no labels for who was entitled to know what, so a permission violation can't even be scored after the fact. And an agent that mishandles a lien deadline doesn't fail loudly. It produces a document that looks fine until a right has quietly expired.
Our answer is a simulator, in the strict sense that pilots mean it: a complete, consequence-faithful environment where the stakes are rehearsed before they're real. Blocks generates entire construction projects (the team, the schedule, every RFI, submittal, change order, daily log, and pay application) deterministically, under real jurisdiction rules researched from primary sources. Because code produced every fact, the correct answer to every question exists by construction. That's what makes honest grading possible, and it's the one thing a demo, a deck, or a reference call can never give you.
Every fact with consequence (dates, amounts, statuses, entitlements) is computed deterministically. Language models render only the human voice of the record. That separation is why an answer key exists at all.
Jurisdiction profiles are researched against statutes and primary sources, independently re-derived by a second pass that doesn't trust the first, and given a final sign-off. "How do you know your answer key is right" has a written, auditable answer.
Scores freeze against the exact versions that produced them. Assessments are never edited to move a number, and improvements land as new versions with an explicit re-baseline, because a quietly edited test makes every prior result meaningless.
We apply our own standard to ourselves: a change to the product's behavior that lowers its evaluation scores does not ship. Assurance buyers ask for exactly this process claim, and we can show it rather than assert it.
Platform-neutral by policy; the public benchmark runs on a submitter model where vendors publish their own numbers. A referee that belongs to one platform, or ambushes vendors with scores, isn't a referee.
We report checkable statements, cited to the rule and the record, with consequences in the project's own units. We don't sell a badge, because a badge is worth the body behind it, and we'd rather earn that standing than assert it.
Blocks is hosted-only and engagement-driven: bake-offs for general contractors making agent procurements, diagnostics and per-release regression for the teams building the agents, and a public benchmark methodology with reference baselines. Two jurisdictions are reference-grade today, Ontario and Texas, with more added as engagements require them, through the same research, verification, and sign-off pipeline.
If you're evaluating construction AI, or building it, we'd like to show you the environment.
Thirty minutes: a live project, a live assessment, and the report your risk team would receive.
Book a demo