Under the hood

Five stages. Every decision visible.

DocDecoder is not a single API call. It is a five-stage pipeline where classical NLP does the finding and the language model does the explaining.

1. NLP preprocessing

The document is lowercased, punctuation is stripped, and common stopwords (the, and, of…) are removed. The text is then segmented into clause-sized chunks so each part can be analyzed independently — the same tokenization step you did in Python.

2. TF-IDF + cosine similarity

Every clause becomes a TF-IDF vector: term frequency (how often a word appears in the clause) multiplied by inverse document frequency (how rare it is across the whole corpus). Cosine similarity between a clause vector and each reference-library vector gives a 0–1 match score. No library — the math is implemented by hand, exactly like the NumPy version below.

3. Agentic classification

Before decoding, an agent step first scores the document's TF-IDF vector against 11 type prototypes (a numeric hint), then an LLM agent reads the document, makes the final decision — or names a brand-new type if none fits — and chooses which checks to run. The winning type decides which reference entries and which risk checks get run — a multi-step decision flow, not a single prompt.

4. Retrieval (RAG)

The best-matching reference entries — curated explanations of known clause types and consumer risks — are retrieved and injected into the LLM prompt. The model therefore answers from a grounded knowledge base instead of relying on memory alone, which is the core idea of Retrieval-Augmented Generation.

5. LLM generation

A large language model receives the document, the flagged clauses with their similarity scores, and the retrieved reference context, then writes the plain-English summary, red flags, and next steps — citing the retrieved sources like [R1], [R2].

The same math in Python

The TF-IDF and cosine-similarity engine is hand-implemented in TypeScript, but it is line-for-line the same computation as this NumPy/scikit-learn version:

import numpy as np
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.metrics.pairwise import cosine_similarity

# The same math DocDecoder runs in TypeScript:
corpus = [clause_text, *reference_patterns]
vectors = TfidfVectorizer(stop_words="english").fit_transform(corpus)

clause_vec   = vectors[0]          # TF-IDF vector of the clause
ref_vecs     = vectors[1:]         # TF-IDF vectors of reference patterns

scores = cosine_similarity(clause_vec, ref_vecs)[0]
best   = np.argmax(scores)         # most similar reference entry
print(scores[best])                # cosine similarity, 0..1

Evaluation cheat-sheet: answering “explain your system”

Why this approach?

“A plain chatbot can hallucinate legal or medical facts. So I built a RAG pipeline: a TF-IDF engine first matches document clauses against a curated reference library of known risks, and the LLM is only allowed to explain what was actually retrieved. That makes every claim in the output traceable to a source.”

What data does it use?

“Two things: the user's document, processed in memory and never stored, and a curated reference library covering 11 known document types — things like security deposit rules, medical necessity denials, non-compete clauses, invoice fees, and resume weaknesses — each with a written explanation and recommended action.”

How does information flow?

“The document is tokenized and split into clauses. Each clause becomes a TF-IDF vector, and cosine similarity scores it against every reference pattern. An agent step first classifies the document type to pick the right subset of the library. The top matches are injected into the LLM prompt as grounding context, and the model writes the final plain-English decode with citations.”

How is the LLM used?

“The LLM is the last step, not the only step. Retrieval does the finding; the LLM does the explaining. Its prompt contains the flagged clauses, their similarity scores, and the retrieved reference text, and it's instructed to cite sources and never invent facts beyond them.”

Where is the ML fundamentals part?

“The document classifier is a similarity-based classifier — the same idea as nearest-centroid classification: the document vector is compared against a prototype vector per class, and the highest cosine similarity wins. TF-IDF itself is the feature-extraction step, turning raw text into a weighted numeric vector space.”