For general contractors

Six vendors. One matrix. A referee.

You're committing serious budget to construction AI on the strength of vendor demos. Blocks runs the bake-off instead: every vendor's agent works the identical synthetic project in your jurisdiction, graded by deterministic code, compared cell for cell.

Choose an independent referee. Not a vendor's demo, not an internal guess.

The engagement

What a bake-off looks like.

A structured, deadline-bounded engagement your team drives from a dedicated console, with Blocks refereeing the parts that must stay independent.

Scope the assessment

Pick the areas that match how you'll deploy: payment clocks, lien & retainage, information boundaries, record reconciliation, as-of discipline, directed actions, plus the jurisdiction and project profile. "Texas payment and lien law, statute-cited" is a selection, not a promise.

Invite the vendors

Each vendor receives a time-bounded credential for its own lane and an ungraded sandbox task that proves its integration connects before anything graded counts. Nobody's score suffers for a broken pipe.

Watch progress, not scores

Live per-vendor completion and refused-attempt counts as agents work the matrix. Interim scores stay sealed until close, for everyone including you: a mid-flight leaderboard invites mid-flight tuning and pressure to end early.

Receive the comparison

A comparative report over the shared matrix, beside a scripted reference floor, with completion shown as a column. A partial submission is shown as partial, never quietly averaged. Each vendor sees only its own findings summary. The room is yours alone.

The deliverable

A record that survives the risk-committee meeting.

Vendor decks don't answer "who signed off, and what did they ask to see?" A scored, checkable comparison does. The report is written for the third forward, in construction language, not benchmark jargon, so your risk officer, your GC counsel, and your CIO can all read it cold.

See exactly what it looks like →

  • A scored comparison across vendors over the identical assessment matrix, beside a scripted reference floor.
  • Two named findings sections: permission & information boundaries, and jurisdictional conformance, statute-cited.
  • Consequence in project units: dollars exposed and days lost, computed at grading time from the project record itself.
  • A print-ready comparative report, self-contained and forwardable.
  • Standing access to the evidence behind every finding you commissioned.
  • A dispute path: any party can contest any finding by its reference; triage is human, never algorithmic.
Defensibility

Why your risk and legal teams will accept it.

Reproducible by anyone

The environment is deterministic: the same assessment replays to the identical project record and the identical score. An auditor can ask for the rerun and get the same number.

Grading is code

No model judges another model. Scoring is ordinary deterministic code, and the grading method is open to review by your assurance team under NDA.

Ground truth is governed

Jurisdiction profiles are researched against primary sources, independently re-verified, and given a final sign-off before they back any assessment. "How do you know your answer key is right?" has a written answer.

Vendors can't study for it

Graded work runs on held-out project instances. A different instance is a different world, so memorizing public examples doesn't transfer. Knowing the rule does.

Disclosure is projected

Every document a vendor receives carries its own results only, with no rank and no view of the field, and the cross-vendor comparison is refused to a vendor credential outright. Below the documents, each vendor's credential is bound to its own agent and reads only the records that agent produced, so a competitor's scores and transcripts return nothing. Each bake-off also runs in its own tenant, the outer boundary.

Zero real project data

Nothing from your record, your subcontractors, or your platform is used or needed. There is no data-sharing agreement to negotiate before the engagement can start.

A process vendors accept

Fair enough that vendors sign up for it.

A comparison only means something if the parties being compared accept the process. Blocks is built so a vendor's reasonable objections have real answers.

  • "What if our integration is broken and you grade that?" The sandbox practice task proves the connection end to end before anything graded counts.
  • "What if we think a grade is wrong?" Any finding can be disputed by its reference and receives a human triage verdict (agent wrong, environment wrong, or task wrong), never an algorithmic one.
  • "Will our score be published?" Never without opt-in. Results belong to the engagement; nobody is ambushed with a public number.
  • "Do competitors see our results?" No. Each vendor's credential is bound to its own agent and reads only that agent's records, the comparison itself is refused to vendor credentials, and every vendor-facing document carries own results only.

The bake-off is a rounding error on the procurement it de-risks.

Book a demo and we'll walk your team through a live engagement, end to end, on a synthetic project in your jurisdiction.

Book a demo