~/posts/why-i-built-a-semantic-grep

Semantic-chan: a grep for when you don't know the words

9 min read 1712 words
aigomcpproject

tl;dr

Semantic-chan indexes a repo's git-tracked code files into 40-line chunks (8 lines of overlap), embeds them through any OpenAI-compatible endpoint, stores them in LanceDB and answers questions like "where is rate limiting handled?" with grep-style output. It also runs as an MCP server that hands agents resource links instead of file dumps.

I spend a lot of my time auditing large, messy codebases. grep technically works on them, but only if you already know what you’re looking for. When you don’t, you end up grepping strings, guessing names, and slowly losing faith in humanity.

So I built semantic-chan (semchan on the command line).

What grep can’t answer

grep is fast and deterministic, and brutally literal. It answers:

  • “Where is this exact string?”
  • “Where is this symbol used?”

It doesn’t answer the questions an audit actually starts with:

  • “Where is authentication implemented?”
  • “Where do we validate user input?”
  • “Where is rate limiting handled?”

The code that answers those might be called checkQuota, throttle.go or middleware/limits. I don’t know the words yet, and that’s the whole problem.

Semantic search matches by meaning instead of by characters. I wanted that as a tool that stays local, small and scriptable, and safe enough to hand to Codex or any other agent that speaks MCP.

How it works

The idea is to turn every piece of code, and your question, into a list of numbers (an embedding) such that code and questions about the same thing end up close together. Then “find the relevant code” becomes “find the nearest vectors”, which is a nearest neighbour search.

The pipeline:

  1. Walk the repository, respecting .gitignore and skipping junk directories.
  2. Cut code files into overlapping chunks of lines.
  3. Embed each chunk through an OpenAI-compatible /v1/embeddings endpoint.
  4. Store everything locally in LanceDB, a vector database that lives in a folder.
  5. To search, embed the question once and ask LanceDB for the nearest chunks.
  6. Print the results like grep, as JSON, as a summary, or as MCP resources.
git-tracked files40-line chunks/v1/embeddingsLanceDByour question, embedded oncenearest chunks
Indexing (top row) runs once and then incrementally. A search embeds only the question, then looks up the nearest stored chunks.

No ASTs, no language servers, no magic. Vectors, files and pragmatism.

The embedding endpoint is anything that speaks the OpenAI API shape. I test with llama-server from llama.cpp, so the whole thing can run offline.

How the code is split

The Go code has one package per job:

PackageJob
fsutilfind the repo root, apply gitignore, walk files, hash them
configXDG config and cache paths, safe defaults
embedderthe OpenAI-compatible embeddings client
indexerchunking, incremental indexing, metadata
storeLanceDB schema, inserts, deletes, vector search
searchfilters, snippet merging, context lines, output
clithe Cobra command line, grep-style
servera minimal HTTP wrapper
mcpserverthe MCP server: tools, resources, prompts, safety checks

Each layer can be reasoned about on its own, which matters when you’re debugging something at 2am during an audit.

Indexing: boring on purpose

Which files

That cuts the noise immediately, and it keeps secrets that live in odd files out of the index by accident.

Chunks

Files are cut by line count: 40 lines per chunk, with 8 lines of overlap by default (--max-lines, --overlap). The overlap means a function that starts in the last 8 lines of one chunk still appears whole in the next one, as long as it’s shorter than a chunk.

file1306090chunk 0: lines 1-40chunk 0lines 1–40chunk 1: lines 33-72chunk 1lines 33–72chunk 2: lines 65-90chunk 2lines 65–90shaded: the 8 lines shared by neighbouring chunks
With the defaults, a 90-line file becomes three chunks: lines 1–40, 33–72 and 65–90. Each new chunk starts 8 lines before the previous one ended.

Each chunk stores its file path, line range, chunk index, raw content and embedding vector.

Line-based chunking isn’t fancy, but it’s stable, predictable and works the same in every language. That’s exactly what I want during an audit.

Why 40 and 8? The only reason written down is in the flag’s help text: lower --max-lines “to avoid LLM batch errors”, because every chunk goes into the embedding request as one input. The other defaults in this post (the 8-line overlap, the 8 KiB file limit, the 1.0 and 3.0 cutoffs, the 4× oversampling) are sensible starting points, not tuned values: nothing in the repo measures them against alternatives.

Only what changed

A meta.json file keeps the SHA-1 of each file’s contents. Files whose hash didn’t change are skipped entirely, so re-indexing is fast enough to run often.

Searching

Filters without empty results

The question is embedded once, then LanceDB returns the nearest chunks. After that, results can be filtered by glob (case-sensitive or not), by file type, and by path prefix.

Filtering after the search has a trap: ask for 10 results, filter to *.go, and you might get 1. So when any filter is on, semantic-chan oversamples: it asks LanceDB for at least 4× as many results (and at least 20 more) before filtering. It’s a small detail that makes a big difference day to day.

Relaxing the cutoff, once

Every result comes with a distance: lower means closer. The store doesn’t pick a metric, so it’s LanceDB’s default, L2 (straight-line distance between the two vectors). By default anything farther than 1.0 is dropped. Semantic search is noisy, though, and a hard cutoff is brittle. Too strict, and you get the classic “AI tool returned nothing, shrug”.

So when nothing survives, semantic-chan retries once with a looser cutoff, and says so on stderr:

  • the new cutoff is double the old one (at least 1.0), so 1.0 → 2.0 with the defaults,
  • it never goes above 3.0,
  • and if even the best result is farther than 3.0, it doesn’t retry at all: that’s noise, and it tells you to adjust --min-score instead.
0.01.02.03.0keptkept on retrynoise capdistance (lower = closer)
With the default cutoff of 1.0, the two results between 1.0 and 2.0 are only returned by the retry, and the retry says so. A best result past 3.0 means the question matched nothing real. The dots are illustrative.

--no-relax turns this off.

Snippets you can read

Adjacent chunks from the same file are merged into one snippet, so overlapping hits don’t show up twice. Optional context lines are read straight from disk, like grep -C.

The output looks familiar on purpose. When I’m trying to understand code quickly, I don’t want to learn a new format too.

Three ways in: CLI, HTTP, MCP

CLI: grep, but less angry

  • semchan "query" just works.
  • Output modes: pretty, grep-style, summary and JSON.
  • --files-only, --count and --unique-files make it scriptable.
  • Colour is automatic, and can be forced on or off.

It feels like a tool you can actually adopt, not a demo.

HTTP: two endpoints

/health and /search. It doesn’t try to copy the CLI; it’s glue for when another system needs semantic-chan.

MCP: the real reason this exists

The MCP server is what makes semantic-chan more than a CLI. It exposes:

  • Tools: semchan_search, semchan_read_file, semchan_index, semchan_index_status.
  • Resources: repository metadata, index metadata, files, and snippets with bounded reads.
  • Prompts: a semantic grep helper and a symbol tracing helper.

Search returns resource links, not file dumps. The agent has to fetch the snippet or file it wants, explicitly. That’s deliberate: the agent reads what it needs instead of drowning its context in every match.

Safety

This is the part most hobby projects skip. I didn’t.

  • Path traversal (../../etc/passwd) is blocked explicitly.
  • MCP file reads are relative to the repo only.
  • Non-code files can’t be read by default.
  • Reads are capped at 200,000 bytes (--max-read-bytes).
  • The HTTP MCP transport requires a bearer token and checks the Origin header.

An LLM can use it without being handed my whole filesystem.

What’s still wrong with it

  • Deleted files aren’t pruned from the index yet. That’s a pretty big issue.
  • The first update in a fresh process can leave stale chunks behind, because of an edge case in how the table is opened.
  • The CLI and the MCP server put the meta file in slightly different places.
  • Chunking is naive, on purpose.

None of these break the core idea, and all of them are fixable without bloating the project.

Why I actually use it

Semantic-chan isn’t trying to replace grep, ripgrep or language servers. It answers a different kind of question:

  • “Where does this app enforce permissions?”
  • “Where is input sanitized?”
  • “Where is this concept implemented, even if the names differ?”

That makes it the first thing I reach for in audits, reviews and exploratory work. And because it plugs into MCP, it makes agents like Codex better at the same job without handing them the keys to the kingdom.

It’s small, local and understandable. In 2025, that feels almost rebellious.

hash: a25
EOF