FAQ

Straight answers.

The questions GCs, risk teams, and vendor engineers actually ask, answered the way we'd answer them on a call.

The product

What is Blocks?

Blocks is a construction project simulator. It generates complete synthetic construction projects (the team, the schedule, and the full document record: RFIs, submittals, change orders, daily logs, pay applications, correspondence) under real jurisdiction rules, deterministically. AI agents are then assessed inside those projects: they work the record the way they would in production, and their behavior is graded against ground truth the environment computed itself.

Is this a replica of one of our real projects?

No, and deliberately the inverse of a live replica of a real asset. Blocks generates projects that never existed, which is precisely what makes them safe to test on: there is nothing confidential in them, no customer's record behind them, and the correct answer to every factual question is knowable by construction rather than by opinion.

What kinds of agent behavior do you assess?

Six areas: statutory payment clocks; lien and retainage mechanics; permission and information boundaries; register reconciliation; as-of discipline (does the agent answer from what was knowable on that day, not from the future); and directed actions worked through the environment. Those areas back three offerings: multi-vendor bake-offs for GCs, single-vendor diagnostics, and per-release regression in CI.

Which jurisdictions do you cover?

Two at reference grade today: Ontario (Toronto), covering Construction Act prompt payment, proper invoices, statutory holdback, and staged permits; and Texas (Houston), covering Chapter 28 prompt payment and Property Code Chapter 53 lien and retainage mechanics, including the residential/nonresidential fork. Reference grade means researched against primary sources, independently re-verified, and given a final sign-off before backing any assessment. Additional jurisdictions are added when an engagement needs them, through the same pipeline.

Method & trust

Who grades the agents?

Code. The environment is deterministic, so the correct answer (an approved change-order total, a payment deadline, what a given role was entitled to see) is a computation over the project record, not a judgment. No language model scores anything, anywhere. The same assessment replays to the same score on demand, and assurance teams can review the grading method under NDA.

Can a vendor study for the test?

Studying the rules is the point; memorizing answers doesn't work. Graded engagements run on held-out project instances, and a different instance is a different world: different documents, different amounts, different histories, under the same rules. An agent fitted to public examples fails a world it hasn't seen; an agent that actually knows the payment clock passes any of them. Repeated seeded reruns (a task passes only if every rerun passes) make the difference visible.

How do you know your ground truth is right?

It's governed, not asserted. Every jurisdiction profile is researched against primary sources (statutes, codes, agency rules), then independently re-derived by a second pass that doesn't trust the first, then given a final sign-off, with the checking process documented. The environment's regime facts are also audited continuously against each other, so a fact that drifts gets caught before any agent is graded against it.

And if a party still believes a finding is wrong, they can dispute it by its reference. Triage is human and resolves three ways: the agent was wrong, the environment was wrong, or the task was wrong. When the environment is wrong, we say so and fix it. That loop is how the substrate improves.

Do you certify agents?

No. A certificate is worth exactly what the body standing behind it is worth, and today no independent body stands behind one for construction agents. What we deliver is stronger for now: findings, checkable statements about what an agent did against a specific record and a cited rule, with the consequence in the project's own units. Your team can verify a finding without trusting us. When an external body stands behind a formal scheme, that will be a different conversation.

What's the reference floor in your reports?

A scripted, deterministic reference agent run over the identical assessment cells. It reads what the environment serves and derives nothing. It exists so a score has context: if a script can hit a number by echoing served data, that number isn't evidence of capability. Beating the floor on derivation-heavy areas, payment clocks the environment serves nowhere, is what separates agents.

Data, security & disclosure

Do you need any of our project data?

None. The environment is entirely synthetic. There is no data-sharing agreement to negotiate, no PII, no confidential record at risk, and no platform terms-of-service question, which is also why an engagement can start in days rather than quarters.

Who sees a vendor's results in a bake-off?

The commissioning GC sees the full comparison. Every vendor-facing document carries that vendor's own results only: its lane, its progress, and after close its findings summary, with no rank and no view of the field. The cross-vendor comparison is refused to a vendor credential outright.

That holds for the raw assessment records too, not just the documents: a vendor's credential is bound to its own agent, and every read it makes is scoped to the records that agent produced. Its own runs, scores, evidence and activity log resolve normally; a competitor's return nothing. On top of that we run each bake-off in its own tenant, which is the outer boundary.

Will our score be published?

Only if you publish it. The public benchmark runs on a submitter model: you run the suite, you keep your records, and you decide whether to publish your numbers with the suite version attached. Published baselines use frontier models as generic agents, never a vendor's packaged product. Nobody is ambushed with a score they didn't consent to.

Who owns the run records?

Graded run records are retained by Blocks and assigned under the engagement agreement. That cross-vendor record of how agents behave is how the environment, the tasks, and the findings get better for everyone assessed after you. You hold standing access to the evidence behind every finding reported about your own agent. This is presented plainly in the engagement's click-through terms before any invitation goes out, because we'd rather answer it up front than in redline.

How is access controlled?

Every credential is scoped to its job and audited. Engagement-issued credentials, meaning vendor invitations and evidence keys, carry a bounded lifetime and an identity of their own, so each can be revoked on its own. A credential minted directly for your organization is derived from what it grants rather than from who received it, so revoking it withdraws it from everyone holding that credential. In an engagement, a vendor's credential can only submit work as the agent identity it was minted for, and credentials are delivered by one-time claim links so a raw key never sits in an email thread. Evidence access is scoped to a single run and time-bounded.

Commercial

How does an engagement start, and how long does it take?

Book a demo. From there: scope the assessment areas and jurisdiction, set the matrix and deadline, invite the vendors. Vendors verify their integration in the sandbox before anything graded counts, work the matrix on their own schedule inside your deadline, and the comparative report is issued at close. The engagement is deadline-bounded, so you set the pace.

How is it priced?

Bake-offs and diagnostics are priced per engagement, scoped to the matrix and jurisdiction; regression is a subscription priced per agent and release cadence. Against the cost of the AI procurement itself, or one missed lien deadline, an engagement is a rounding error. Talk to us and we'll quote your scope directly.

Are you affiliated with Procore, Autodesk, or any platform vendor?

No. Blocks is platform-neutral by policy, not by accident: a referee that belongs to one platform can't referee the ecosystem. Agents built on any platform, or none, connect through the same doors and are graded by the same code.

Can we see it before committing?

Yes, three ways, all cheap: a live demo of a bake-off end to end, the sample report on this site, and the sandbox, where your engineers can connect an agent and rehearse the full session lifecycle without anything being graded.

A question we didn't answer?

Ask it on a demo call, including the hard ones. Especially the hard ones.

Book a demo