besthvacaeo.agency Directory

Methodology

What these scores are, what they are not, and how you would reach the same numbers yourself.

What this leaderboard measures, and what it does not

No data source available to us can independently verify how often any agency gets its clients named by an answer engine, or how that changed over time. Not one. The tools that track this at scale are licensed products, and even those see what a set of prompts returned rather than what a homeowner in a particular town actually saw.

So this site does not score performance, because it cannot. It scores what each company discloses and how it says it measures, both of which any reader can check from published pages in an afternoon. No number on this site is a measured result about an agency's outcomes, and a directory that tells you otherwise about a third party should be asked where the data came from.

On a site whose argument is that most of this field cannot be falsified, saying where our own evidence stops is the position rather than an apology for it.

The five criteria

Each is scored 0 to 5 from published material, then weighted. The weights are below with the reasoning for each, including where a weight was changed and why.

Weights sum to 100. Each criterion links to the full test it is scored against.
Criterion Weight What earns a high score
Named engine coverage 25% How many consumer answer engines are named specifically, and whether they are treated as distinct systems rather than as “AI”. Using a retired name such as SGE caps this at 2.
Measurement method shown 20% Is there a stated prompt set size, a run cadence, a geography, and a visible reporting artefact? “We track AI mentions” with nothing behind it scores 2 at most.
Citation change over time 10% A named client, a stated method, and a before-and-after across dates. Sample and illustrative dashboards score 2. Nothing published scores 0.
Entity and schema competence 30% Specific schema types named, entity consistency work described, knowledge-graph and third-party identity work, crawler access. Generic “we do schema” scores 2.
Falsifiability 15% Does the agency state what failure would look like, disclaim what it cannot control, or publish a limit on its own claims? Unsourced statistics and unsupported mechanism claims score 0 and are named in the profile.

Why each weight sits where it does

Named engine coverage · 25%
Checkable from published material in a minute, and it tells a buyer what they are actually purchasing. Raised from 20 because it is one of only two criteria in this set that a reader can fully verify without trusting us.
Measurement method shown · 20%
Held at 20. Measuring what an engine says is demonstrable and repeatable; proving why it said it is not. Those are opposite things, and only the second belongs in the unknowable column.
Citation change over time · 10%
Cut from 20. Every company in the set scores between 0 and 2, so at the old weight it was mostly dead weight in the arithmetic rather than a real separator. It stays on the board because the day someone publishes a real time series, it should count.
Entity and schema competence · 30%
Raised from 20, and it now carries the most weight. This is the layer with a mechanism you can inspect: markup validates or it throws errors, and crawler access is a line in a file. It is the demonstrable half of the argument on our anchor page, so weighting it heavily makes the leaderboard agree with the rest of the site.
Falsifiability · 15%
Held at 15. It cannot be raised while measurement method is held down, because a claim is falsifiable only where a stated method exists by which it could be shown false. Falsifiability at 15 with method near zero would score whether an agency writes modest sentences rather than whether it would ever know it was wrong.

The reserved sixth criterion

If a source becomes available that can track what engines actually answered, per model, over months, a sixth criterion for measured citation performance will be added. Two rules for that are published now, before there is any incentive to bend them.

First, the five weights above will renormalise proportionally against a stated share rather than being renegotiated one by one. Second, any performance score will appear only alongside the named source that produced it and the date of the run. Writing that down while it costs nothing is the only time it is worth anything.

Why the list is short

Because the overlap between companies that measure this properly and companies that serve HVAC contractors is thin. Of the four in our set that show a real measurement method, three are software vendors rather than agencies, and only one is a home-services agency in the ordinary sense. The companies doing this rigorously mostly do not work with contractors, and the companies working with contractors mostly do not measure it.

There is a second reason the list is shorter than it looks, which is that we refuse to score from a search result. Three companies appear on this site with no score at all because their pages could not be read. One of them has a search snippet suggesting it would place near the top. It is still unscored, because a snippet is not a page.

What moves a company up or down

Up: publishing a prompt set, naming engines individually rather than as “AI”, publishing a dated client result with the method attached, and stating what would count as the work having failed. Every one of those is a decision to disclose rather than a capability to acquire, which is why scores here can change quickly.

Down: unsourced statistics presented as fact, retired terminology left on a live page, and mechanism claims with no evidence behind them. Where we find one, it is named on the company's scorecard with the reason, so you can judge whether we were fair.

Sources, dates and review

Every scorecard carries the date its page was read. Live search results were read from the United States in English on that date. Scores are reviewed quarterly and the next review is 12 December 2026. Every change lands in the corrections log with its date.

Challenging a score

Point at the page that proves us wrong. Scores come from published material, so a page already stating a prompt set, a cadence or a failure condition will move a score as soon as we have read it. We will not remove a listing on request and we will not raise a score in exchange for anything. We will re-read and publish the result either way, in the corrections log, with the date.

All scorecards Corrections log The argument behind these criteria