Answer-pipeline
attribution.
We tested two ways of preparing company evidence for an answer, then used mathematics from Precise to measure what each change contributed and how they interacted. Measurements were collected July 17 and attribution was replayed September 6, 2026.
Konstant using Precise · Internal product evaluation
Measured July 17 · Attribution replayed September 6, 2026
Improve how evidence reaches an answer
Mathematical methods for measuring what parts contribute and how they interact
Leave this filter out of the combination
A better answer
needs the right evidence.
Gary brings company knowledge into answers. Konstant’s product team was testing two changes to that pipeline: a filter that chooses which extracted records to keep, and a compact format that puts each selected record into one row while removing repeated surrounding structure.
Each change improved exact answer coverage by itself in this comparison. The team also measured them together, using the same source support, questions and answer requirements.
Konstant supplied those measured outcomes to Precise’s existing attribution method. Precise received a numerical table; the attribution call did not need the underlying company documents.
Measured exact
answer coverage.
The compact format carried 95 of 178 required points exactly into the answers. Adding the filter reduced that to 80. The evaluator judged each required point as exact, partial or missed; this table shows exact coverage.
| Configuration | Exactly covered | Coverage | Evidence bytes |
|---|---|---|---|
| Baseline | 78 / 178 | 43.8% | 562,590 |
| Filter only | 84 / 178 | 47.2% | 545,528 |
| Compact format only | 95 / 178 | 53.4% | 354,091 |
| Both changes | 80 / 178 | 44.9% | 339,086 |
Two source documents, one development comparison. “Evidence bytes” measures the representation supplied to synthesis; it is not a dollar cost or a latency measurement.
A part’s value depends
on what it joins.
The filter’s standalone gain was +3.37 percentage points. Its effect alongside the compact format was −8.43 points. A score based only on the standalone result would miss that reversal.
- Filter contribution
- −2.53 percentage points, averaged across the two orders in which these changes can be added.
- Compact-format contribution
- +3.65 percentage points on the same basis.
- Interaction
- −11.80 percentage points: how far the combined result falls below the sum of the two standalone gains.
Precise computed exact Shapley contributions and the pair interaction from all four measured configurations. The table already identifies the highest-scoring configuration. Attribution adds a consistent allocation of the measured gain and quantifies the interaction; this experiment does not establish that the method found a decision the comparison alone would miss.
The result had
somewhere to go.
Konstant retained the compact format without the filter as a candidate for replication. Work on preserving who said and owned each claim was a separate development concern.
The follow-up replication has a completed run record. The first comparison did not authorize production promotion. Answer quality, attribution accuracy and consistency across further examples still mattered.
Our agents connected Precise’s existing mathematical capability to Konstant’s answer evaluation. This put that expertise to use in a different company’s product and let Konstant build on an existing implementation.
Inspect what went in.
Inspect what came back.
The September run replayed Konstant’s existing adapter locally against the actual Precise implementation. It reproduced the recorded attribution for answer coverage, partial-credit quality and evidence size. It did not regenerate the July answers.
Aggregate evaluation results are available here. Source documents and individual judgments remain private. The evidence record describes how the exact candidate receipts were verified despite a subsequently regenerated parent run manifest.
What each company
actually did.
Konstant defined the answer requirements, generated and evaluated the configurations, supplied the complete measurements and made the development decision. Its adapter called Precise’s separately maintained implementation.
Precise computed each change’s contribution and the interaction between the changes. Those calculations use the measured results from all four configurations.
Boombox supplies reusable operating infrastructure. A separate authenticated Precise research workload ran in Precise’s cloud and retained its result. That hosted run is distinct from this attribution calculation.
Arranger connects results to subsequent work. Its executed next-experiment study calls another Precise method with illustrative probabilities. The study is a separate integration; its inputs and result are available for inspection.