AI / RESEARCH

IVI Defect Triage

Does symptom normalization improve duplicate detection?

A controlled ablation study on whether normalizing noisy defect reports into one-sentence symptom summaries improves embedding-based duplicate detection, extended into a full triage pipeline and a tool-use agent.

Inspired by my internship experience; built in my own time on fully synthetic data, without any internal data, code or processes.

Role

Solo · Study design, data, pipeline, evaluation

Status

● Finished · public repository

Stack

Python · Anthropic API (tool use) · Voyage AI · NumPy

The question.

Does normalizing noisy defect reports into one-sentence symptom summaries improve embedding-based duplicate detection?

Defect reports of this kind can contain a lot of noise unrelated to the defect (bench names, software versions, test case IDs, colleague names). My hypothesis was that this shared noise inflates vector similarity between unrelated tickets. Two reports written on the same bench, against the same software version, look alike to an embedding model even when they describe completely different defects.

The approach.

Each report is reduced to one German sentence describing only the symptom and the affected function: max 20 words, no bench, software version, date or names. Retrieval runs on these summaries; retrieval on the raw report text is the control.

Embeddings are asymmetric voyage-3: corpus documents are embedded with input_type=document, incoming tickets with input_type=query. Similarity is cosine.

The dataset is 40 fully synthetic German tickets generated with a fixed seed. 5 components (Navigation, Voice Assistant, Connectivity, Media, Display), 8 tickets each, severities S1 to S4, split into 20 corpus and 20 test tickets. 8 true duplicate pairs are split across corpus and test set and written with different wording, length and interaction path; five of them mix English technical terms into one side. 4 hard negative pairs come from the same component and describe different defects. All texts contain deliberate noise.

What the retrieval study shows.

ModeRecall@3OverlapMedian margin
summary8/80.0230.256
raw report8/80.1120.126

Finding 01 · Recall@3 is saturated

With 20 corpus tickets, the top 3 already cover 15% of the corpus. Both modes reach 8/8, so the metric does not discriminate between them.

Finding 02 · Separability is the real difference

In neither mode does an absolute score threshold separate duplicates from non-duplicates. But the overlap between the two score ranges is about five times smaller with summaries.

SCORE OVERLAP // DUPLICATES VS NON-DUPLICATESNON-DUPLICATESDUPLICATESRAW TEXT0.112SUMMARY0.023MEDIAN TOP1-TOP2 MARGIN 0.126 → 0.256

Finding 03 · The margin carries the signal

A margin criterion (top1 minus top2, threshold about 0.18) detects 7/8 duplicates with 0/12 false alarms on summaries, versus 2/8 on raw text. With a true duplicate there is exactly one very similar ticket; with a new defect, candidate scores sit close together.

From study to pipeline.

The retrieval finding becomes the middle of a triage flow with four fixed steps:

  • classify: one Claude call returns component, severity and symptom summary.
  • retrieve: Voyage embeddings on the summaries, cosine search.
  • judge: a statistical margin threshold as pre-filter, then an LLM check of the top candidate.
  • create_ticket: a local CSV row plus a notification log line.

Classification on the 20 test tickets: component accuracy 18/20, severity exact 10/20, within one level 18/20. Both misclassifications sit on rule boundaries.

Gold summaries overstate the result.

The retrieval study uses hand-written summaries. With automatically generated summaries, the median duplicate margin drops from 0.256 to 0.178.

The consistency of hand-written summaries is a property of the annotation, not of the method. The retrieval study is therefore an upper bound, not the performance to expect from the running system.

Calibrating the right stage.

Three ways to set the margin threshold of the judge step:

VariantThresholdPrecisionRecallF1
D1 carried over0.181.000.380.55
D2 calibrated for zero false alarms0.201.000.380.55
D3 calibrated for F1 plus LLM filter0.081.000.750.86

In a two-stage design the statistical stage should collect candidates broadly and the LLM stage should filter. A false alarm costs the tester one look; a missed duplicate creates a duplicate ticket.

Fixed pipeline versus tool-use agent.

MetricPipelineAgent
Component accuracy90%95%
Severity accuracy (exact)50%55%
Duplicate precision1.001.00
Duplicate recall0.381.00
Duplicate F10.551.00
API calls per ticket1.153.90
Cost per ticket (USD)0.0028 (estimated)0.0509 (measured)

Pipeline duplicate rows use the zero-false-alarm threshold 0.20. At the D3 operating point the pipeline reaches precision 1.00, recall 0.75, F1 0.86 at 1.5 API calls per ticket.

The agent has three tools: search_similar_tickets, get_ticket_details and create_ticket, with a maximum of 8 calls per ticket and a median of 4. When it is unsure, it reformulates the search or looks up the full text, which lets it find the two pairs every threshold misses.

The price is more than 10x the cost per ticket (0.051 USD vs. under 0.005 USD for the pipeline) and decisions that are less reproducible.

Three mistakes I corrected.

Mistake 01 · Symmetric instead of asymmetric similarity

I first judged separability on document-to-document similarity, which suggested full separability. The real query-to-document path scores 0.15 to 0.20 lower. Lesson: evaluate on the same path the system uses.

Mistake 02 · Label leakage in an agent tool

search_similar_tickets initially returned component and severity of the candidates, exposing gold labels. After removing them, agent component accuracy dropped from 100% to 95%. Only the 95% is comparable.

Mistake 03 · Unsuitable calibration criterion

Zero false alarms at the statistical stage optimizes the wrong stage and gives away recall (0.38 versus 0.75).

Limitations.

  • Self-made data. I created the data and the gold annotation, so the study measures whether retrieval reproduces the construction intent, not performance on real tickets.
  • Text length. The texts (60 to 160 words) are longer than real tickets on purpose.
  • Small sample. 8 positive pairs give wide confidence intervals: the Clopper-Pearson lower bound is about 0.63 at 8/8 and about 0.47 at 7/8.
  • Not a deployed system. No UI, no real ticketing integration.

Code

github.com/Haichennn/symptom-normalization-retrieval

Full reproduction steps in the README.

Stack

Python · Anthropic API (tool use) · Voyage AI · NumPy

Haichen Duan