POSITION-CONTROLLED EVALUATION SUITE

v1.2 strict source-span survival

Normalized contiguous required-span containment is the post-hoc primary metric over frozen outputs. Every annotated source span must survive as one contiguous normalized passage; prohibited content and budget violations still fail the case. No compressor or QA calls were rerun.

Normalized contiguous required-span containment by output-token budget.
Evaluated method or adapter1282565121024
Trimwise Lexical52.5%60.0%61.9%66.2%
Trimwise Hybrid49.4%62.5%66.9%69.4%
RECOMP NQ extractive sentence adapter27.5%30.6%35.0%35.0%
LLMLingua GPT-2 token-pruning adapter3.1%7.5%13.8%22.5%
LongLLMLingua GPT-2 single-context adapter0.6%4.4%6.9%16.2%

Evidence survival and operational feasibility

At 256 tokens, span survival ignores prohibited text and budget compliance; feasible case pass requires all three. This separates retrieval failure from an output that violates the shared evaluator budget.

Evaluated method or adapterSpan survivalFeasible case passBudget violation
Trimwise Lexical71.9%60.0%0.0%
Trimwise Hybrid74.4%62.5%0.0%
RECOMP NQ extractive sentence adapter30.6%30.6%0.0%
LLMLingua GPT-2 token-pruning adapter15.0%7.5%43.8%
LongLLMLingua GPT-2 single-context adapter15.6%4.4%69.4%

Local ordered-retention sensitivities

These 256-token values are feasible case pass under the predeclared post-hoc metrics. The 80% and 90% tests require ordered matching within one reference-length output window; complete contiguous containment remains primary.

Evaluated method or adapterContiguousLocal ordered 80%Local ordered 90%
Trimwise Lexical60.0%61.3%61.3%
Trimwise Hybrid62.5%63.7%63.7%
RECOMP NQ extractive sentence adapter30.6%32.5%31.9%
LLMLingua GPT-2 token-pruning adapter7.5%8.1%8.1%
LongLLMLingua GPT-2 single-context adapter4.4%5.0%5.0%

Natural, relocated, and self-source sensitivity checks

The natural-only subset is unbalanced because it contains 15 natural end cases. The 25 relocated rows are a construction diagnostic, not a causal substitute for natural placement. The self-source check removes every real-trimwise-* case.

Strict case pass separately for natural placement and controlled relocation cohorts.
Evaluated method or adapterNatural n=135Relocated n=25Without self-source n=150
Trimwise Lexical62.2%48.0%60.0%
Trimwise Hybrid63.0%60.0%63.3%
RECOMP NQ extractive sentence adapter26.7%52.0%32.0%
LLMLingua GPT-2 token-pruning adapter0.0%48.0%8.0%
LongLLMLingua GPT-2 single-context adapter0.0%28.0%4.7%

Where the strict result succeeds and fails

All values below are 256-token feasible strict case pass. Evidence positions are balanced in the 160-case primary suite; task counts are not equal.

Evaluated method or adapterBeginningMiddleEndMultiple
Trimwise Lexical72.5%65.0%57.5%45.0%
Trimwise Hybrid77.5%65.0%62.5%45.0%
RECOMP NQ extractive sentence adapter35.0%40.0%32.5%15.0%
LLMLingua GPT-2 token-pruning adapter0.0%0.0%30.0%0.0%
LongLLMLingua GPT-2 single-context adapter0.0%0.0%17.5%0.0%
Evaluated method or adapterAdversarialEvidence QAInstructionProcedureReal sourceStructured
Trimwise Lexical35.3%82.6%0.0%70.0%63.3%90.5%
Trimwise Hybrid35.3%80.4%0.0%70.0%73.3%100.0%
RECOMP NQ extractive sentence adapter35.3%4.3%11.5%70.0%36.7%61.9%
LLMLingua GPT-2 token-pruning adapter11.8%0.0%7.7%10.0%20.0%0.0%
LongLLMLingua GPT-2 single-context adapter0.0%0.0%3.8%5.0%16.7%0.0%

Paired uncertainty: Hybrid versus RECOMP adapter

Fixed-seed percentile bootstrap intervals over paired case-level differences in the 160-case frozen suite. They are descriptive uncertainty intervals for the post-hoc v1.2 analysis, not preregistered hypothesis tests.

BudgetHybrid - RECOMPFixed-seed 95% paired bootstrap interval
128+21.9 pp[+13.1, +31.2] pp
256+31.9 pp[+22.5, +41.2] pp
512+31.9 pp[+21.9, +41.9] pp
1,024+34.4 pp[+25.0, +43.8] pp

The complete local-80%, local-90%, and both Trimwise configuration intervals are in the paired- bootstrap CSV. See the benchmark documentation for the protocol, frozen manifest, complete CSVs, and reproduction commands.