Field Notes · 6 MIN
How do you build an attribution engine and prove it is accurate?
Field notes on a wallet attribution engine. Rate limits, serial jobs, and an eval that reports right, wrong, and no answer separately.
You build it as a pipeline with an evaluation harness from day one, and you report accuracy as three separate numbers: right, wrong, and no answer. An engine that guesses on every input looks accurate until it is wrong in front of someone who acts on it. This post covers how we structured one attribution engine, how we ran the eval under real rate limits, and what we refuse to claim from a small sample.
- Report right, wrong, and no answer separately. A confident wrong answer and an honest abstention are different outcomes with different costs.
- A small test set gives a number about that set only. Never quote it as a general accuracy rate.
- Keyless public APIs have hard rate limits, so two heavy jobs on the same IP will starve each other.
- Stream results per case and run long evals in batches, so one failure does not lose the whole run.
- Nactore is an AI-native software engineering partner that ships to production with evals, and this is the shape those evals take.
What is an attribution engine?
An attribution engine takes a messy input and returns the most likely real-world entity behind it, with supporting evidence. Ours took a suspect crypto wallet address and returned the nearest exchange, the deposit address involved, a draft of the request to send, and a PDF of the evidence. The problem comes from a public-sector challenge, and the output is meant to be acted on by an investigator.
That last clause drives every design choice. When a person acts on the output, a wrong answer costs far more than no answer. So the engine has to be able to say "I could not determine this," and the evaluation has to reward that honestly.
Why report right, wrong, and no answer separately?
A single accuracy percentage hides the thing that matters. Consider two engines that both score the same on "correct." One abstains on the cases it cannot solve. The other guesses on all of them. They are not equivalent, and only a three-way split shows the difference.
| Outcome | What it means | Cost to the user |
|---|---|---|
| Right | Correct entity, with evidence | Saves time |
| No answer | Engine abstained | Costs a manual look |
| Wrong | Confident but incorrect | Can mislead a real decision |
We track wrong as its own number and treat driving it down as the primary goal. Raising right at the cost of more wrong is usually a bad trade. This is the same discipline we describe in how to evaluate LLM output quality, applied to a non-LLM pipeline.
How should you handle a small evaluation set?
Honestly. A small labeled set gives you a number about that set. It does not give you a general accuracy rate. Our rule is that we never quote the small-sample rate as if it were general, and we say the sample size next to the number every time.
Small sets are still worth building. They catch regressions, expose failure categories, and show where the engine abstains. They just cannot support a claim like "the engine is 90 percent accurate in the wild." Say what the set measures, and stop there.
How do you run an eval against rate-limited APIs?
Our engine read chain data from a public API that we called without an API key. At the time we worked on it, the keyless tier allowed roughly one request per second per IP, and a breach triggered a short suspension. Those limits are the vendor's to change, so check the current documentation before you depend on a number.
What we learned from living under that limit is general.
- Never run two heavy jobs at once. An eval and a live trace on the same IP will breach the limit together and both will fail or stall.
- Handle 429 explicitly. When the API says slow down, wait longer than the suspension window, then retry.
- Stream results per case. Write each case's result as it completes, so a crash at case 40 of 50 does not erase cases 1 to 39.
- Run in batches. We ran batches of about sixteen cases, because our background job tooling had a time limit on a single run.
- Resume, do not restart. A batch runner that skips finished cases turns a failure into a minor delay.
Treat the rate limit as part of the system design, not a nuisance. If you need more throughput, the answer is a keyed tier or a different data source, not parallelism against a hard cap.
How do you keep the engine honest as it grows?
Attribution engines tend to improve by adding heuristics, and each heuristic can raise right while quietly raising wrong. Three habits keep that visible.
- Re-run the full eval on every change. A heuristic that helps one case may break two others.
- Diff the per-case results, not just the totals. Totals can stay flat while individual cases swap from right to wrong.
- Keep abstention as a first-class outcome. Every new rule should say what it does when it is unsure.
For LLM-powered pipelines the same habits apply, and we cover the setup in AI evals before production.
Write the "no answer" path before you write the clever path. If abstaining is awkward to implement, the engine will guess, and the eval will show it as wrong answers.
Why do screenshots and UI review matter for an engine?
The output of an attribution engine is a report a human reads, so the interface is part of the product. Two process notes from this build.
First, we stopped delegating browser screenshots to subagents after they stalled repeatedly. We now generate screenshots ourselves with a small script that intercepts routes to serve fixtures, then hand the images to a design reviewer. That is faster and deterministic.
Second, every interface change still goes through a design review gate before it ships. A reviewer looking at real screenshots at desktop and phone widths catches problems that code review misses. See from prototype to production: an AI checklist for where that gate sits in our release flow.
Frequently asked questions
What is the difference between "wrong" and "no answer"?
A wrong answer is confident and incorrect. A no-answer is an honest abstention. The first can mislead a real decision, while the second just sends the user to a manual check.
Can a small test set prove the engine works?
It can show how the engine behaves on that set and catch regressions. It cannot prove a general accuracy rate, so do not quote it as one.
Why not run the eval in parallel to finish faster?
Under a per-IP rate limit, parallel jobs breach the limit together and stall. Stream results per case and run batches serially instead.
Do these methods apply to LLM features?
Yes. Separating right, wrong, and abstain, streaming per-case results, and re-running the eval on every change all carry over to model-based features.
Want this built for your team? Book a free 30-minute call.
Want to apply this to your business?
Book a free 30-minute call. We will tell you what we would do first.