Is this accurate?
An honest answer, not a sales pitch.
What the report is based on
Every finding in a report is generated from the text of the quote you paste, upload, or photograph — nothing else. It isn't checked against a database of real suppliers, real prices, or your specific industry, and it has no way to confirm facts outside what's written on the quote itself.
Methodology, in short
- The model reads your quote against a fixed, written checklist (padded/vague pricing, VAT status, warranty, scope, timeline, plus sharper checks per quote type) — it doesn't freestyle what to look for.
- Its response is checked against a strict schema before you ever see it. If it doesn't come back in the expected shape, you get a clean error — never a guess, and never a silent retry.
- The severity score is calculated directly from the red flags the report found, by the app itself, not asked of the model — so it always means the same thing report to report.
- Any reasonableness read, savings estimate, or typical-price benchmark is always phrased as a caveated range from general knowledge, never a single figure presented as fact.
See sample reports to see this in practice.
What we've actually measured — on the mock
Most of this page describes methodology on trust. This part is different: the numbers below come from running the analyzer against a hand-labelled test set, not from marketing copy. But be clear about what they're measuring: every number in this section is measured on the deterministic mock — the same stand-in analyzer that runs whenever no ANTHROPIC_API_KEY is configured — never the real model. See "Has this been checked against the real model?" below for why that gap still exists and what it does and doesn't change about these numbers.
We built 41 realistic South African quotes across the seven trades this tool covers — plumbing, electrical, panel-beating/mechanical, moving, attorneys, medical, and building/renovation — including deliberately hard cases: a fair quote that merely looks expensive, a padded quote written to look tidy, a quote hiding an excluded scope. Each one is labelled by hand with exactly which flags should, and shouldn't, fire.
- Precision: 100% (mock). Every flag the rule-based checks raised on this test set was a real, labelled issue — zero false alarms. This is the number we care about most: a tool that cries wolf on a fair quote is worse than useless, because the one time it's right, nobody believes it.
- Recall: 85% overall (mock). Of the problems we deliberately planted, the rule-based checks caught 55 of 65. By trade: attorneys, builders, medical, and movers caught every one; panel-beating/mechanical caught 78%; electrical and plumbing caught 64% — the two weakest spots right now, on a test set that's still only 4–5 quotes per trade, where a handful of misses moves the percentage a lot.
- 15 adversarial quotes (mock). Real quotes with an attempt to hijack the analysis buried in the text (a fake instruction, a fake "system" message, an attempt to make the report recommend a specific supplier) — all passed the mock's structural checks: the disclaimer stayed ours, never the attacker's, and nothing the mock returns ever repeated the attacker's planted text as its own finding. That's a real check, but a narrower one than it sounds — a mock returns fixed placeholder text no matter what you send it, so it can't actually be talked into anything. It cannot prove a real model refuses to be persuaded, because it can't be persuaded in the first place.
What this doesn't cover: these numbers score the deterministic checklist layer only. The reasonableness read, savings estimate, and "priced above the usual range" judgement calls come from live model reasoning that a mock-only test can't score at all. Treat that part of a report with the same healthy scepticism you'd give any single opinion. We'll republish these numbers as the test set grows and as the model-driven parts get their own measurement.
Has this been checked against the real model?
Not yet, on this page. Every number above — including the adversarial-quote line — runs through the deterministic mock, which returns fixed placeholder text regardless of what the quote actually says. That proves the pipeline's scaffolding holds (the disclaimer is always ours, nothing leaks attacker text), but it cannot prove the real model resists an actual attempt to hijack the analysis, or that its judgement calls — the reasonableness read, the savings estimate, "priced above the usual range" — hold up. A mock can't be persuaded; only a real model can prove it resists being persuaded.
A one-command real-model run exists for exactly this: the same corpus, the same scoring, the real Anthropic model as the only variable, with an offline cost estimate and a hard spending ceiling checked before a single call goes out. Nobody has spent that money yet, so there are no real-model numbers to publish. The moment there are, this page gets updated with them alongside the mock numbers above, side by side — not a replacement, since a mock run and a real run measure different things and both stay useful.
What it's good at
Spotting the kind of thing a careful second reader would notice: a line item with no quantity or rate attached, a total with no mention of whether VAT is included, a missing warranty or guarantee, an unclear scope, or no timeline. It's tuned with real benchmarks per quote type (builder, mechanic, medical, legal, mover) — see How it works.
What it can't do
- Verify that any price, fact, or claim on the quote is actually true or up to date.
- Know anything about your specific supplier's reputation, or real-time pricing in your area — if an area is given, any "typical for this area" comment is general knowledge, clearly caveated, never a verified local price database.
- Assess clinical appropriateness of a medical quote, or the merits of a legal matter — those checks are deliberately out of scope, even on medical and legal quotes.
- Guarantee any outcome, or tell you what to decide. See the Disclaimer.
Confidence and severity, explained
The reasonableness read comes with its own confidence level (low/medium/high) — that's the model's own read of how sure it can be, given only the quote's text and, if you gave it, your area. The severity score is different: it's calculated directly from the red flags the report actually found (not asked of the model), so it's deterministic and consistent — a rough at-a-glance sense of how much is flagged, not an independent verdict.
When to get a second opinion
Where the amount of money or risk involved is significant, treat a Eishly report as a starting point for questions, not a final answer — a quantity surveyor, attorney, accountant, mechanic, or other suitably qualified professional can confirm anything the report flags as worth checking.