← All projects
04 · Agent evaluation

PRAMAAN — Auditing What the Evaluation Actually Measures

Checking Whether I Knew What I Thought I Knew

Solution video
Five minutes — the problem, one full execution, and the finding
Watch on YouTube →

What this project became

It started as a claim checker for advice circulating about stocks listed on NEPSE. It ended as a study of a narrower and more awkward question: how do you know your metric measures what you think it measures?

The score I was watching was already at its maximum before I started. A baseline with no documents at all scored 3 out of 3, and five runs later the shipped agent scored 3 out of 3. There was never any room above it to move into — and every time it did move, a measure I had not thought to record moved the other way.

Every time I checked whether I knew something, I found I did not — so the next thing I built was another way of checking. The repository records that drift in its own proportions.

What it is Files Lines
The subject under test agent.py, baseline.py 384
The apparatus that measures it validate.py, test_validate.py, check_results.py, check_divergence.py 548
The record of what the measurements meant CHANGELOG.md, DECISIONS.md, CASES.md, REPRODUCE.md, README.md 1,611

Counted with wc -l at commit c28f8e3. Seven result files, twenty stored model responses, none overwritten or hand-edited.

More code measuring the system than composing it — and more documentation than either, at four times the agent’s line count. The claim checker is the subject on the bench, not the product.

The six questions

Here is the answer, and it is uncomfortable: take away the thing you believe the number is measuring, and see whether the number notices. If it does not notice, it was never measuring that.

I did it six times — three by taking something away, three by turning the same suspicion on the instruments doing the measuring. Every one came back against me. These six are the transferable part of this page, and everything below them is the evidence behind each answer.

The question What it returned when I asked it
Null Does my metric separate my system from a stripped-down version of it? Mine didn’t. A baseline with no documents at all scored full marks, 3/3 — the same as the shipped agent, four iterations later, because full marks was the top of the scale.
Contamination Does my evaluation data already contain the answer I am testing for? Mine did, for a whole iteration. I had written my own conclusions into the source files the agent reads, so it was retrieving answers I had planted.
Ablation Have I run it once with my favourite component switched off? Not until after I had shipped it. When I did, the loop turned out to cost 3.4× the time and 1.7× the tokens, and to move nothing I was scoring.
Instrument Is the checker I trust actually tested? Not until I wrote 23 tests over it — which found four defects that reading the code had not.
Divergence Is the checker driving my loop the same function as the one that scores me? Mine was — called with different text. The runtime check was stricter in four of eight attempts, and never looser — so every violation it reacted to was invisible to the score.
Drift Has my ground truth moved since I started — and did I record it? Mine moved. One verdict was revised after two runs had completed. The change and its cost are dated in DECISIONS.md.

Findings from three cases do not generalise. Questions do — which is why this list is made of them, and why each answer beside it is a cost I paid rather than a law I am claiming.

They divide in two, and the division is the practical part. The first three each cost a run. Score a stripped-down version. Take your own conclusions out of the data and score it again. Switch a component off and score it again. Each needs the model, so each has a bill attached — and each can come back against you.

The last three cost nothing at all. test_validate.py runs with no API key, no network and no dependencies. check_divergence.py reads results already sitting on disk and writes nothing. Asking whether the ground truth moved is a diff against git history. They are the three cheapest checks in the project, they need no permission and no budget, and two of them I did not run until the last day.

Ask them in the order above NCAIDD — it is the order I asked them in, and the order that spends least before it tells you something. Shuffled, they spell a word that is easier to carry than a list: CANDID. Contamination · Ablation · Null are the three that need a run of the model. Divergence · Instrument · Drift need nothing but a text editor and the files you already have. CAN costs. DID is free.

Note: NCAIDD is the questioning flow; CANDID is the mnemonic that makes it easy to remember.

How the parts fit together
One rule set, two callers — one audit over the rules, one over the callers
claim + fixed corpus nothing is retrieved from the web agent.py builds the prompt, drives the loop the model one call per attempt prompt JSON results/*.json every attempt, with its own prompt and its own token count validate.py the score every attempt reads validate_response() status is one of three · the quote is in its file every number in the figure is in that quote sources AS WRITTEN line breaks intact violations → retry, max 3 sources NORMALISED whitespace collapsed test_validate.py 23 tests — is the rule itself correct? check_divergence.py what the loop saw at the time, beside what the scorer sees tests the rule itself checks the two paths agree THE RUN THE RECORD THE RULE THE AUDIT
Both callers run the same three rules. Only the text they hand them differs — and that is the finding.

How to read it

  1. Top centre. A claim and a fixed set of documents go in. Nothing is retrieved from the web, so any run can be re-scored later from its own recorded prompt.
  2. agent.py and the model. One prompt out, one JSON response back. That pair is a single attempt.
  3. Down to the record. Every attempt is written to results/*.json with its own prompt and its own token count. Nothing is overwritten — which is what makes the divergence audit at the bottom possible at all.
  4. The arrow down the middle — agent.py hands the response to validate_response() together with the source files as written, line breaks intact. If a rule fails, the arrow beside it carries the violations back up and the model tries again — up to three attempts.
  5. The arrow entering from the right — once the run is finished, validate.py reads the results and calls the same function. But it hands it the sources with the whitespace collapsed.
  6. The bottom row. test_validate.py asks whether the rule is correct at all. check_divergence.py asks whether those two arrows were handing the rule the same thing. They were not.

Each instrument exists because the previous one returned “you do not actually know that.” Read the diagram top to bottom and it is the same question, asked one level deeper each time.

The subject under test

This is built for one kind of NEPSE investor: someone in a full-time job who invests on the side, without the hours — or in some cases the financial background — to check whether the advice reaching them rests on any evidence. And the claims that reach them are rarely false. They are true figures with the context stripped off.

Test claim · synthetic

“Nabil declared 30% dividend this year. Bank FD is giving only 4%. Why keep money in fixed deposit when a blue chip bank is paying 30%? Long term hold, guaranteed income.”

Nabil did declare 30.00% — for fiscal year 2078/79, three years earlier. For the year the claim is about, 2081/82, it declared 12.50% and no bonus. And a dividend percentage is calculated on the Rs 100 par value, not on the Rs 540.20 the share costs. So Rs 12.50 ÷ Rs 540.20 = 2.31%, against ordinary individual fixed deposit rates of 2.75% to 4.55% across the three banks in the sources.

What the claim actually compares
The dividend it calls 30%, against the deposit rates it calls 4%
0% 1% 2% 3% 4% 5% Nabil Bank · normal 2.80–4.55% NIC Asia Bank 2.75–4.00% Everest Bank 2.75–4.05% Nabil dividend yield 2.31%
Rs 12.50 ÷ Rs 540.20 = 2.31% — below the floor of every deposit range in the sources, which is the opposite of what the claim asserts.

Every number is true. They support the opposite conclusion. Checking it by hand takes four lookups, and the step that decides it — par value, not market price — is the one most people skip.

PRAMAAN — evidence — takes a claim, breaks it into separate assertions, and rules each one verified, contradicted or not_found against a fixed corpus. Every figure must arrive with the exact quote it came from and the file that quote lives in. If the sources do not settle an assertion, it must say not_found rather than estimate. It never advises buying or selling.

How agentic is this, honestly?

Only the validation stage acts, checks its own output and acts again — the minimum that qualifies. No planning, no tool use, no memory — and question three asks whether even that much earns its cost.

1 · Does my metric separate my system from a stripped-down version of it?

baseline.py — one API call, one line of instruction, no documents at all.

What it got right

  • Reached the correct verdict on all three cases
  • Knew a dividend is quoted on par value, unprompted
  • Rejected “guaranteed income” on its own

What it invented

“you might have to pay a Market Price of NPR 600 or more”

“your actual return on investment (yield) might only be 3% to 5%”

Both fabricated. The real figure, 2.31%, sits below the floor of that range.

Correct answers, invented evidence, 0/3 figures traceable to a source. And a verdict score of 3/3 — identical to the shipped agent, four iterations later. That is the ablation, and it arrived on day one: whatever my metric was rewarding, grounding was not it.

The deeper problem is not that the two scores tied. It is that they could not have done anything else. Three out of three is the top of the scale, and the stripped-down version was already there — so the measurement had no room left to record anything I built afterwards. A metric that is full before you start is not a hard metric to beat. It is a metric that cannot be beaten, and reading a tie as a disappointing result rather than a broken instrument cost me three iterations.

2 · Does my evaluation data already contain the answer I am testing for?

Iteration 1 looked like a clean win. Figures traceable to a source went 0/3 to 3/3, and the attribution error went from missed to caught, 0/1 to 1/1. Both numbers moved in the right direction and stayed there.

Then I read my own source files properly. Of the six substantive quotes Iteration 1 returned, exactly one was a published figure. Here they are, verbatim from results/agent_v1.json:

The quote the agent returned What it really was
Cash dividend: 12.50% published A figure from the company’s own record.
Range across all three banks: 2.75% to 4.55% for individual normal fixed deposits. mine A range synthesised from three separate rate cards — not a published figure.
Nabil’s dividend yield of 2.31% is below the entire deposit range, and the share carries price risk a deposit does not. mine My conclusion, with the arithmetic already done in the file.
…the share carries price risk a deposit does not. mine The same sentence again, cited for a second assertion.
That figure belongs to the sector, not to EBL, and covers a different period… mine The attribution answer, written by me, under a heading naming the check.
Price to book: 729.50 / 246.74 = 2.96x — trading at nearly three times book value mine Arithmetic already worked out in the file, not derived by the agent.

One published figure. The other five were lines already sitting in the file — analysis, not data. What looked like the agent reasoning was mostly the agent reading my own file back to me.

The worst instance

Case 3’s attribution finding — the result I was most pleased with — quoted a sentence I had put into sources/ebl.md myself, under a heading that named the check being performed. The agent had not detected anything. It had retrieved my answer.

Iteration 2 was the fix: every line of my own analysis moved out of sources/ and into notes/, a folder the agent never sees. Only published figures remained, plus one neutral line per file saying what it does and does not contain. Then I ran it again.

The attribution finding survived unaided. The agent found the sector line on its own, and derived its own supporting arithmetic — price to book 729.50 / 246.74 = 2.96 — rather than copying a number I had already worked out for it.

It cost a verdict. Agreement fell from 3/3 to 2/3, because Case 1 came back partly supported: the agent marked “Bank FD is giving only 4%” verified against a single rate-card row, and never computed the yield that decides the case.

The decision

Report both. A 3 of 3 built on a corpus that contained the answers is worth less than a 2 of 3 that does not.

That is the run that made every later number mean something.

3 · Have I run it once with my favourite component switched off?

The correction loop is the thing I had built that whole iteration around. A control run with it disabled scores identically on both machine-checked metrics.

Loop on Loop off
Verdicts 3 / 3 3 / 3
Schema violations scored 0 0
Mean time 70.1s 20.7s
Mean tokens 10,741 6,373

The loop cost 3.4× the time and 1.7× the tokens and moved the scored metric by nothing. It did repair four violations of the rule it enforces, which the control leaves standing. Against the rule that scores this project, it changed nothing. That is a negative result and it is reported as one. Both runs are in the repository: results/agent_v4.json and results/agent_v5.json.

So the loop is not what the scored metric was responding to. The change that mattered was decontaminating the sources — question two above — and not because it raised a score. It lowered one.

4 · Is the checker I trust actually tested?

Three rules, with no model in the loop. The quote rule is checked against the run’s own recorded prompt, not today’s files, so retrospective scoring stays honest after the corpus changes.

Rule What it checks What it cannot catch
1 · Status The status is one of exactly three permitted strings. Nothing — this one is complete for what it claims.
2 · Quote The quote appears in the file it names. Whether the quote has anything to do with the assertion beside it.
3 · Figure Every number in the reported figure appears in its own quote. Whether that number means what the assertion says it means.

Rules 2 and 3 establish that a citation is well-formed. Neither can establish that it is apt — which is the limit no mechanical rule closes.

test_validate.py runs twenty-three tests: twelve over the three rules — that each catches what it should and passes what it should — seven over the whitespace, source-parsing and response-shape helpers those rules depend on, and four that pin known gaps in the rules, written to fail if a rule is ever tightened.

What the tests found

Four defects, none of which reading the code had shown me — and one of them was living in a regular expression I had read several times over. A checker you have only read is one you are trusting, not testing.

No API key, no network, no dependencies, and it finishes in well under a second. The gaps were documented rather than fixed, because every run was scored under the current rules and three of them can no longer be re-run.

One of those gaps could have hollowed out the headline result: rule 2 only inspects quotes that exist, so a run could in principle reach zero violations by omitting quotes rather than by getting them right. Every assertion in all four agent runs was inspected before that gap was accepted. Every verified and every contradicted assertion carries a quote, and Iteration 4 quotes even its not_found assertions, where none is owed. The gap is real and unexercised — recorded in DECISIONS.md, 30 August.

5 · Is the checker driving my loop the same function as the one that scores me?

Found on the last day of the build, while writing the video script. The loop and the scorer call the same validator with different text: agent.py passes the source files as written, line breaks intact; validate.py passes them whitespace-normalised. A quote spanning a line break fails the first and passes the second.

Nobody chose that rule. It falls out of one line — the quote check inside the validator:

elif normalise(quote) not in sources[named_file]:

The quote is flattened inside the function. The source is not — it arrives however the caller passed it, and the docstring asks only for “its full text”, never in what form. agent.py hands over the file as read. validate.py hands over the same file with its whitespace collapsed. Both satisfy the contract as written. Only one of them can be the rule.

check_divergence.py re-scores every stored attempt the way a finished run is scored, and prints that beside the violation count the loop recorded live at the time. Eight attempts have per-attempt records. Four agree. In all four that disagree, the runtime check saw violations and the scorer sees zero — stricter in one direction, never the other. The control is in that count: with the loop disabled the check still ran and still flagged, it just never corrected.

Violations detected live Repaired? Violations when scored
Iteration 4 — loop on 4 Yes, at 70.1s and 10,741 mean tokens 0
Control v5 — loop off 4 No — it shipped with them on its record 0

The control tied, but not because the loop was redundant. It tied because the scoring rule never charged for the thing the loop was repairing.

One execution, start to finish

The run where that divergence surfaced. Case 3 · EBL undervaluation · results/agent_v4.json.

“EBL net profit up 32% this quarter, best in the sector. Still trading below book value, market has not priced it in yet. Undervalued blue chip, accumulate now.”

Step What happened Cost
Attempt 1 The agent finds the 32% — and finds it belongs to the whole commercial banking sector, not to EBL. Marks the assertion contradicted and quotes the line. 6,417 tokens
Validator Rejects the response. Not the verdict — the citation.
assertion 1: quote does not appear in sources/ebl.md
1 violation
Attempt 2 The same response is returned with the violation named and an instruction to fix that and change nothing else. Comes back clean. 5,758 tokens

What changed between the two attempts

The quote the agent gave
Attempt 1 — rejected "Commercial banks' combined net profit rose 32.33% to NPR 69.78 arba in Q4 2082/83."
Attempt 2 — accepted "Commercial banks' combined net profit rose 32.33% to NPR 69.78 arba in"
sources/ebl.md
lines 80 and 81
Commercial banks' combined net profit rose 32.33% to NPR 69.78 arba in ↵
Q4 2082/83. Nabil Bank is reported as leading the sector.

Same verdict. Same figures. Same summary, byte for byte. The only difference is that the quote now stops at the line break — and under the rule that scores this project, the rejected version was already valid. The loop threw out a good citation and the replacement it forced carries less: the sector attribution survives, the reporting period does not.

The finding

The model was not learning to cite better. It was learning where the file wraps.

The instruction asked for “the exact line from the source document” and told the model to change nothing else. It also said what to do when a quote came up short: “If one quoted line cannot support the whole figure, quote the line that can, or reduce the figure to what the quote supports.”

So the truncation is not the model working around my check. It is the model following my instruction.

Read as a check rather than a defect, this is the baseline test arriving from the other direction. A newline changed the validator’s answer while nothing about the evidence changed — and a check that moves for reasons unrelated to the thing it claims to measure is not measuring that thing.

6 · Has my ground truth moved since I started — and did I record it?

Yes — once. Case 1’s verdict in CASES.md was revised from partly supported to unsupported — after both the baseline and Iteration 1 had already run.

What moved it. The claim quotes a 30% dividend. I had accepted the figure as genuine and judged the claim misleading only in how it framed a true number. During a file audit I checked the source: Nabil declared 12.50% cash and no bonus for FY 2081/2082. The quoted figure is contradicted by the source for the year the claim is about, so the verdict became unsupported.

Why the timing does not invalidate the measurement

A ground truth adjusted after seeing results is worthless unless you can say what moved it. What moved this one was the source document, not either system’s output.

Both result files that existed at that point were re-scored against the corrected verdict, and both moved identically — baseline and agent each went from 2/3 to 3/3. That is confirmed rather than assumed: both files record “unsupported” on Case 1. The comparison between the two systems is unaffected.

The disclosure sits in REPRODUCE.md, so anyone reproducing these numbers meets it before they meet the scores.

What it cost

A finding. I had recorded a third weakness in the baseline — that it condemns a whole claim rather than separating the accurate part from the misleading part — and built it on the assumption that the 30% was real for the current year. Once the figure was checked, the gap dissolved.

I removed it rather than leave a conclusion standing on a premise I never verified. Two weaknesses, not three: it verifies nothing, and it never asks whose figure a number is.

A figure taken on trust and never traced to its source — the exact failure this project was built to catch, committed by me.

The final comparison

Five runs plus a control, scored on the same three cases. Schema violations are scored mechanically by validate.py. Verdict agreement is against hand-scored verdicts — and Case 1’s was revised from partly supported to unsupported after the baseline and Iteration 1 had already run, because the source contradicted my original reading. That revision, and what it did to the measurements, is dated in DECISIONS.md.

Run What changed Verdicts Schema violations Mean time Mean tokens
Baseline No documents at all 3 / 3 — 15.8s 1,832
Iteration 1 Grounding + structured output 3 / 3 5 16.4s 5,338
Iteration 2 My conclusions stripped from sources 2 / 3 2 15.3s 4,805
Iteration 3 Range and ratio rules in the instruction 3 / 3 4 20.3s 5,620
Iteration 4 Validation loop — shipped 3 / 3 0 70.1s 10,741
Control (v5) Same instruction, loop disabled 3 / 3 0 20.7s 6,373

A schema violation is a single assertion breaking one of the validator's three rules: a status outside the three permitted values, a quote that does not exist in the file it names, or a figure whose numbers are absent from its own quote. The verdict can still be right — it just cannot be checked. One response can break several.

The metric that was already full Biggest change from the baseline Biggest change across the iterations
Verdict agreement: 3, 3, 2, 3, 3. Three out of three is full marks, and the baseline — with no documents at all — had it from the start. Nothing built after that could show up here. Figures traceable to a source: 0/3 → 3/3, from Iteration 1. Attribution error: 0/1 → 1/1 — but unaided only from Iteration 2, once I stopped planting the answer in the corpus. Schema violations: 5 → 0. Not a baseline comparison — free text has no schema to break. And the control scores 0 with the loop off, so the loop is not what closed them. What did is unmeasured: Iteration 4 added two instruction rules alongside the loop, and both runs carry them.

The axis I was not watching

Two measures, shown separately because they are different scales — and the baseline appears on only one of them.

Verdict agreement
Correct verdicts out of three cases
0 1 2 3 3 Baseline 3 Iter 1 2 Iter 2 3 Iter 3 3 Iter 4 3 Control
Verdict agreement, out of a maximum of 3, reads 3, 3, 2, 3, 3 across the baseline and four iterations, and 3 for the control. The baseline with no documents is already at the maximum, so it matches the shipped agent.
Schema violations
Lower is better — the baseline has no bar, because free text has no schema to break
0 1 2 3 4 5 5 Iter 1 2 Iter 2 4 Iter 3 0 Iter 4 0 Control
Schema violations run 5, 2, 4, 0 across the iterations, and 0 for the control run.

Of the three iterations that predate the validator, Iteration 2 is my worst on verdicts and my best on citations — and Iteration 3, which fixed a real reasoning defect, doubled the violation count, 2 to 4. Across Iterations 1 to 4 there are eleven violations in total, and ten of them are citation failures. python validate.py lists them.

The changelog

One variable per iteration — except the last, which is why the control run exists.

What I changed What it taught me
Baseline What a user gets today: paste the claim into a chat window. Reasons well. Verifies nothing.
Iteration 1 Gave it the published figures; required a quote and file per figure. Traceability 0/3 → 3/3. But I had written my own conclusions into the sources — so the attribution catch measured retrieval, not detection.
Iteration 2 Stripped my analysis out of the sources into notes/. Only published figures remain. The finding survived unaided, and the agent derived its own arithmetic instead of copying mine. Cost one verdict.
Iteration 3 Two rules added: state the full range when comparing; compute any ratio the claim turns on. 3/3 restored and earned. Citations got twice as bad.
Iteration 4 Built validate.py and wired it into the agent: fail → return the violations → retry — and restated two of the validator’s own three rules in the instruction, so the model is told the standard before being corrected against it. Two variables in one iteration: the control isolates the loop and leaves what those two rules contributed unmeasured. Schema violations to zero. 3.5× the time and 1.9× the tokens of Iteration 3 — and 3.4× / 1.7× against the control.

What this does not establish

Stated here because the project’s entire argument is that unchecked claims are the problem.

  • Three evaluation cases · one attribution case. A case study, not statistical power. The attribution result rests on a single case — the weakest evidence here.
  • The reference is my own judgement. Verdicts were hand-scored by me and checked against no second reader. Verdict agreement therefore measures agreement with one person, not with an independently established answer.
  • Findings are local, the method is what transfers. One model, one corpus, no repeated runs and no variance estimate. Nothing here supports a general claim about correction loops; it supports a claim about this one, measured this way, with the control that shows it.
  • Five defects shipped unfixed. Four are validator gaps, left because tightening a rule on the last day would re-score five completed runs against a contract the agent was never given. The fifth is the runtime / scorer divergence in question five, left for a different reason: aligning the two paths on shipping day would leave every recorded run describing a system that no longer exists. The reasoning for each is dated in DECISIONS.md.
  • No user has used it. It ships as an evaluation harness, not a product. The README says so plainly.
  • The scored metric was saturated before the first iteration. The baseline scored full marks, so verdict agreement had no room left to record anything built after it. It was the wrong thing to optimise, and I only know that because I measured it and published the result.

And the limit no mechanical rule closes: the validator proves a citation is well-formed. It cannot tell whether it is apt. In Iteration 3 the agent marked “Bonus shares provide free shares to investors.” as contradicted and cited | 2081/2082 | 0.00% | 0.00% | No dividend | — traceable, correctly structured, and addressing nothing.

Reproduce it

Four scripts · no API key · no network:

Script What it reproduces
check_results.py timings, tokens, verdicts
validate.py schema violations per run
test_validate.py 23 tests over the validator
check_divergence.py the runtime / scorer divergence

19 dated decision records, two carrying predictions registered before the run · 7 result files, never hand-edited or overwritten · a failed run kept because a failed run is evidence too · full git history in the repository.

REPRODUCE.md walks through it end to end. The baseline and the shipped agent cost 8 API calls and about four and a half minutes on a free-tier Gemini key — which is capped at 20 requests per day, so the guide is written around that limit.

Hot take

The axis you aren’t measuring is the one your agent is moving on.

The score I was watching was verdict agreement. Baseline 3/3, then 3/3, 2/3, 3/3. Three iterations in, by that number, I had built nothing — and that number could not have told me otherwise. It was full when I started.

The fix is not “measure more things” — you cannot act on that. It is narrower, and it is uncomfortable: remove the part you’re proudest of, and run it again.

Take your documents away and see whether your score survives. Mine didn’t. And when the thing doing the measuring is code you wrote, point the same test at that.

Why “measure more things” is not the fix

It does not tell you which thing. The list of properties you could score is endless, and nothing on it announces itself as the one that matters. I only thought to score whether a citation was well-formed because I had already built a validator that defined it. “Measure more” would never have pointed me there — it is an aspiration, and an aspiration has no next action attached to it.

A new metric can be exactly as invalid as the old one. Adding measurements multiplies observations, not validity. The metric I added is a case in point: the same function that computed it was, when called from the other side of this system, enforcing a stricter rule — so the number I trusted and the number the loop acted on came apart on half the recorded attempts.

Addition and removal answer different questions. Another metric can show you that something moved. Only a run without the component can show you that the component was ever what moved it. “Did this contribute?” is a counterfactual question — it asks about the system without the part — and no amount of extra scoring produces a counterfactual. That is why the three sharpest answers above came from taking something away — the documents, the conclusions I had planted, and the correction loop itself — and why the other three came from turning the same suspicion on the instruments doing the measuring.

And removal is a test, not an intention. The control run cost three API calls, and the answer came back against me: the loop I had built that whole iteration around moved nothing I was scoring. “Measure more things” has no such answer waiting for it. Whatever a new metric shows, it cannot show that measuring more was the wrong move — there is always another one to add. Three API calls contradicted me.

Where the work went

Five of the six questions above are not about the agent. They are about the metric, the data behind it, the checker, the second caller of that checker, and the reference I was scoring against. Only the third asks anything about the system itself — whether one part of it earns its cost.

None of those checks was planned. Each one began with me trying to explain something and failing.

  • Reading my own output by hand. The contamination surfaced while I was going through results/agent_v1.json line by line — checking the result I was most pleased with, not one I distrusted.
  • Noticing what had just become load-bearing. I wrote tests for the validator not because it looked wrong, but because Iteration 4 had made it three things in one evening: the source of every schema-violation number, the re-scorer of the three earlier runs, and the gate the model had to pass.
  • Explaining one run to someone else. The largest defect here turned up while I was writing the video script, went to narrate a single execution step by step, and could not account for why a citation had been rejected.

None of those three costs anything. Between them they found the contamination, the validator’s four defects and the divergence — and not one of those came from reading the code.

I did not plan that ratio. It is where the work went once I started checking whether I knew what I thought I knew — and it is why 548 lines of this repository measure the 384 that do the work.

Dipendra Limbu | Nepal | Data Analyst | Business Intelligence | dklimbuz@hotmail.com

All projects LinkedIn → GitHub →