Home / Method
How we grade
The rubric,
in the open.
No black box. Here's exactly how a Houtos grade is built — the axes, their weights, the pipeline, and how we keep measured results honest and separate from previews.
Scan. Score. Synthesize. Recommend.
A website audit methodology is only worth trusting if it runs the same way every time. Ours is a four-stage pipeline — each stage has one job and hands a clean, labeled result to the next.
On this page
The pipeline — Scan, Score, Synthesize, Recommend The 5 axes & 26 metrics — plus the Reputation lens Measured vs Estimated — our honesty promise Why evaluation-led development Who we are & why trust this Frequently asked questionsScan
Pull real signals off the live site with Lighthouse, axe-core, headers and TLS. Nothing copied or stored.
Score
Weighted, confidence-aware model → grades across 5 axes, 26 metrics, each with its own confidence.
Synthesize
Grades become one plain-language read of strengths, risks and trade-offs — confidence labels preserved.
Recommend
A prioritized fix list — biggest impact first, weighted by effort, with expected score lift.
Nothing in this pipeline is a matter of opinion. Two people running the same site through it under the same conditions get the same result. That repeatability is the whole point: it is what turns "this site feels slow" into a number you can defend in a board meeting. Anything observed in the scan is a measured fact; where a real measurement cannot be taken, the gap is recorded — never quietly filled — and only ever backfilled with a clearly labeled Estimated signal.
Five axes, 26 metrics, and a sixth lens.
The website quality score is not one opinion stretched across a page. It is 26 distinct metrics grouped into five axes, each scored on its own and then weighted into the overall grade.
Performance
How fast it actually loads and responds — Core Web Vitals and real load behavior, measured with Lighthouse, not guessed at.
Accessibility (ADA / WCAG)
Real WCAG violations flagged by axe-core and the business and legal risk they carry — explained without acronyms.
SEO & Findability
Whether search engines — and AI answer engines — can actually read, index and cite the site, from structure to metadata.
Security & Privacy
TLS, response headers, and trackers — the gaps attackers and regulators notice first, read directly from the live site.
Trust & Compliance
Credibility signals and the obligations a serious site is expected to meet — framed as business risk, not as legal advice.
The sixth lens — Reputation
Beyond the five scored axes sits a sixth perspective: Reputation — what the wider web says about the brand, from reviews to sentiment. It is treated as a lens rather than a hard-scored axis because much of it cannot be measured directly off your site. Everything in this lens is labeled by confidence, and where it rests on estimate rather than observation, it is marked Estimated so it is never mistaken for a verified reading.
We label our confidence. Always.
Every category carries a confidence tag. We would rather tell you a number is an estimate than dress up a guess as a fact.
A direct reading from a live tool (Lighthouse, axe-core, header inspection). Reproducible, and reported as fact.
Estimated from observable signals while full live measurement rolls out. Clearly flagged, never blended into measured scores.
We are direct about a structural limit: the Grader you can run today still estimates some signals rather than scanning them live. That is exactly why the labeling exists. As the real-scan backend (Lighthouse and axe-core behind a live /api/scan) rolls out, more of every grade moves from Estimated to Measured — and you will see that shift on the page, not hidden behind it.
Our promise is simple: we will never present an estimate as a verified scan. If a number is estimated, it says so.
Please note
This assessment does not constitute legal advice.
Why evaluation-led development.
Most dev shops guess at quality, then ask you to take their word for it. We invert that. We measure quality first — the same way, every time — and let the evidence set the agenda. The grade is not marketing; it is the brief. That is what "evaluation-led development" means, and it is why the methodology lives on the homepage instead of in a drawer. We don't guess. We grade.
Built by people who stress-test their own grade.
Houtos Labs is an evaluation-led web, app and web3 development firm, founded by Ron Clarkson, who also leads Peregrine Precision Systems — the lineage the grading engine ("Peregrine Precision Evaluator") is named for. The Grader is our flagship product and our calling card: a live website-audit tool that proves the claim we make to every client.
We hold ourselves to the same bar we publish. Accuracy is the top internal priority, the engine is actively stress-tested against real sites, and plain-language explanation is a hard requirement on every category — not a nice-to-have. When the deterministic model reaches its ceiling, we say so and build the real scanner rather than dress up the number. The grading methodology, the engine design, and the build roadmap are all maintained as living documents, and the code itself is the system of record.
In short: this page describes how the grade is made because we want it judged on its method, not its confidence. If you find a way to break it, tell us — that is how it gets sharper.
Questions about the method.
How accurate is the website grade?
Accuracy depends on how each signal was gathered, which is exactly why every metric is labeled Measured or Estimated. Measured signals — speed, accessibility violations, security headers — come from running real tools against the live site and are as accurate as those tools. Estimated signals are confidence-aware estimates we use only when a measurement is unavailable, and they are never reported as verified facts. We surface the limits of any estimate honestly rather than inflating a number.
What's the difference between a measured and an estimated signal?
A measured signal is one we directly observed by running a real scan of your live site — for example, a WCAG violation flagged by axe-core or a Largest Contentful Paint timing from Lighthouse. An estimated signal is a confidence-aware value produced from available structural evidence when a direct measurement could not be taken. Measured signals carry a Measured chip; estimated signals carry an Estimated chip. We never present an estimated figure as a verified scan.
What tools do you use?
For measured performance we use Google Lighthouse, the industry-standard performance and best-practices auditor. For measured accessibility we use axe-core, the open-source WCAG rules engine used widely across the accessibility industry. We also read security response headers, TLS configuration and page structure directly. Where a real measurement is not yet available, we fall back to a confidence-aware estimate and label it as such.
Can I trust the score?
You can trust it precisely because it tells you what it does and does not know. The methodology is repeatable: the same site graded twice under the same conditions produces the same result, so the number is defensible rather than a matter of taste. Every figure is traceable to its evidence and labeled by confidence, and we will always tell you when a result is an estimate rather than a verified scan.
How often should I re-grade my website?
Re-grade after any meaningful change — a redesign, a platform migration, new third-party scripts, or content added at scale — and otherwise on a regular cadence such as monthly or quarterly. Because the methodology is repeatable, comparing grades over time is a reliable way to confirm that a fix actually moved the metric it targeted and that nothing regressed.