#graphlens
Граф кода одной командой: ставим graphlens-mcp в проект и перестаём жечь токены на grep В первых двух статьях я сдела...

#graphlens #MCP #граф #кода #кодовые #агенты #Claude #Code #LLM #статический #анализ

Origin | Interest | Match
Граф кода одной командой: ставим graphlens-mcp в проект и перестаём жечь токены на grep
Это третья часть серии. В части 1 я разобрал движок graphlens — что он делает и как устроен. В части 2 я прогнал бенчмарк на 936 запусков и честно показал, где граф окупается, а где нет. Эта статья —...
habr.com
June 28, 2026 at 3:21 AM
How much does context cost an AI coding agent? grep vs graph vs LSP, measured across 936 runs
In my last post I described **graphlens** — what it does, how it works — and along the way I casually claimed that an agent "burns tokens grepping around a repo." I gave exactly **zero** numbers to back that up. This post fixes that. Here are the measurements, the data, and a reproducible harness. Spoiler: the conclusion is not the one I expected going in, and that's the interesting part. ## TL;DR I took **one** agent (Claude Code), changed **exactly one thing** — which MCP server feeds it code context — and ran it over 26 tasks on `apache/superset`. Four "arms": `filesystem` (grep + read), `graphlens` (structural graph), `serena` (LSP), and `codegraph`. Three models (haiku / sonnet / opus), three seeds — **936 runs**. The headline: **the answer flips depending on the kind of task.** * On simple "where is X defined / what does X inherit from" lookups, all four tools are **tied on accuracy**. The only difference is cost (~3×). graphlens is unremarkable here. * On "estimate the blast radius / find every override / disambiguate an overloaded name" tasks, the tools **separate hard** : grep collapses (0.71 accuracy, only 83% of runs even finish, and the ones that do cost **6–24× more**), while the structural tools stay cheap and accurate. If I'd only measured the easy tasks, I'd have written "you don't need a graph, grep is fine." If only the hard ones, "you don't need grep, get a graph." The truth sits in the middle, and it's about **what work you hand the agent.** ## The business case we're actually measuring Picture a familiar situation. You have a large project: hundreds of thousands of lines, a Python backend, a TypeScript front end, legacy code you're scared to touch. You wire an AI agent into it — for review, refactoring, answering questions like "what breaks if I change this method's signature?" The agent can't see the whole repo at once. Something has to feed it context: which functions live where, who calls whom, what inherits from what. And here's an **architectural decision with a price tag** : what exactly do you feed it? There are basically four classes of answer: * **Hand it grep + read** — let it search by text and open files. Zero infrastructure, works everywhere. * **Build a structural code graph** (graphlens) — entity nodes, typed edges, exact answers to "who calls this." * **Stand up an LSP** (serena over a language server) — what your IDE already runs on. * **Use an off-the-shelf code-graph product** (codegraph). Each option costs money (tokens), time (latency), and risk (the agent gives up and hits a turn cap). `apache/superset` is an almost perfect stand-in for this case: ~400k LOC, Python + TypeScript, an `/api/v1/...` boundary between front and back. A big polyglot project — exactly when this question is worth asking. So how much does each option cost? Let's measure. ## Experiment design: change one variable The whole methodology rests on one principle: **fix everything except one thing.** Model, system prompt, settings, task set — constants. Only the context-providing MCP server changes. Then any difference in the numbers is the contribution of that tool, not a config accident. No tool is designated "the baseline to beat." All four are measured on equal footing, and the numbers rank them. ### The four arms Arm | Context provider (MCP server) | Indexing step ---|---|--- `filesystem` | `@modelcontextprotocol/server-filesystem` (read_file + grep) | none `graphlens` | graphlens graph over MCP | `graphlens analyze` `serena` | Serena (LSP) | LSP workspace warm-up `codegraph` | a graph-based competitor | `codegraph init` One detail that matters for fairness: **Claude Code's built-in tools (Read / Grep / Bash, etc.) are disabled.** If you don't take them away, the agent ignores the MCP server and falls back to its usual path — and you'd be measuring the wrong thing. So the harness runs `claude -p` in a clean room: a fresh `CLAUDE_CONFIG_DIR` with only subscription credentials (no hooks, plugins, skills, memory), `--strict-mcp-config` (only this arm's server is visible), `--disallowedTools` on every built-in (an explicit _deny_ , because in headless mode an allow-list alone forbids nothing), and `--allowedTools mcp__<server>` to auto-approve the one server. ### The second axis: models In parallel I varied the model answering the question: Key | model id ---|--- `haiku` | `claude-haiku-4-5` `sonnet` | `claude-sonnet-4-6` `opus` | `claude-opus-4-8` Why a second axis becomes clear near the end: **the optimal tool depends on which model you picked.** That's probably the least obvious finding in the whole thing. Total: 4 arms × 3 models × 26 tasks × 3 seeds = **936 runs** (on Claude Code 2.1.187). ## What counts as an honest measurement Benchmarks are easy to bend toward the conclusion you want. So the rules are fixed up front — without them the numbers aren't trustworthy: * **Gold answers are hand-verified** against source at tag `6.0.0` (every task carries a `file:line` reference). Crucially, **gold is not generated by any tool under test** (not ty, not pyright, not graphlens itself) — otherwise the comparison is biased toward whoever's output you labelled with. Set-task gold is checked with an independent oracle: Python's `ast`. * **The "naive" arm has hands.** `filesystem` is grep + read, not "an agent with no tools." Naive ≠ toolless. * **Index cost is measured separately, once.** grep pays nothing to index; a graph amortizes. You can't mix those currencies. * **There is no determinism.** `temperature=0` does not make these models deterministic. So 3 seeds, and the report shows **the median, not the mean.** * **Versions are recorded** — models and every MCP server — plus a price snapshot and date. * **`cost_usd` is an API-equivalent, not your bill.** The subscription is flat-rate, so `cost_usd` (emitted by the CLI) is what the same tokens _would_ cost via the API. It's **not** your actual invoice, but it is a **correct relative $/task metric** for comparing arms. * **Use a tool, or it doesn't count.** The system prompt forbids answering from memory; a run with zero tool calls is retried (and a stubborn refusal is tagged `__NO_TOOLS__`). Answering "from memory" about a well-known repo wouldn't measure the context provider. And separately: **failure counts as accuracy 0.** If grep hits the 50-turn cap and never produces an answer, that's not "no data" — it's "the tool didn't get there within budget." That's how it's scored. ## The tasks: two regimes, and why you can't blend them 26 tasks split into two classes. **SIMPLE — 20 pinpoint lookups** ("where is X defined / what does X inherit from"). One-point answers, checked by substring: Kind | # | What it probes ---|---|--- `where_defined` | 7 | Python class → defining file `inherits_from` | 5 | Python class → base class `abstract_methods` | 1 | ABC → its abstract methods `ts_where_defined` | 1 | TS hook → defining file `ts_route_call` | 4 | `/api/v1/...` route → the TS hook that calls it `xlang_link` | 2 | TS consumer → Python handler across the API boundary **HARD — 6 blast-radius and disambiguation tasks.** This is the regime where structure and semantics _should_ beat text search — and which pinpoint lookups simply can't measure: Kind | # | What it probes | Scoring ---|---|---|--- `disambiguate` | 2 | an ambiguous bare method name (e.g. `cache_key`, defined on many classes) → _the_ right class | substring `overrides_count` | 2 | the full set of subclasses overriding a base method | **set F1** `impact_set` | 2 | every file calling a given method (the blast radius) | **set F1** Set tasks are scored by F1: reward for recall (find them all), penalty for precision (text search loves to dump every occurrence of `.get_indexes(`). Gold sets are kept small (3–5 elements, one ≈17) so they can be exhaustively checked by hand. ### Why I stratify instead of averaging The set is **deliberately unbalanced** — 20 simple vs 6 hard. A single blended average would be entirely dictated by the easy tasks and would **hide** exactly the difference the hard ones expose. So I report each regime **separately, and never mix them.** And no, I deliberately don't "balance to 50/50" by dropping simple tasks. That would throw away data and statistical power, and open the door to cherry-picking. Stratification neutralizes the skew **without discarding data**. (General principle: if regimes give different answers, it's more honest to show both than to bury the conflict under an average.) ## Results ### SIMPLE — 20 pinpoint lookups Tool | accuracy | complete | tokens | calls | $/task | sec ---|---|---|---|---|---|--- filesystem | 0.97 | 100% | 1780 | 10 | $0.063 | 43 graphlens | 0.98 | 100% | 690 | 3 | $0.038 | 13 serena | 0.99 | 100% | 402 | 3 | $0.031 | 20 codegraph | 0.99 | 100% | 372 | 1 | $0.022 | 10 Accuracy is a **tie** (formally: Friedman χ²=0.40, not significant). The tools differ only on cost — a ~3× spread — and the terse ones win. **graphlens is unremarkable here** — a solid mid-pack. This is exactly the story a benchmark that _only_ measured pinpoint lookups would tell: "structural tools are nice, but grep nearly keeps up, and codegraph gives the cheapest answer." And it would be an **incomplete** truth. ### HARD — 6 blast-radius and disambiguation tasks Tool | accuracy | complete | tokens | calls | $/task | sec ---|---|---|---|---|---|--- filesystem | 0.71 | 83% | 12596 | 27 | $0.424 | 165 graphlens | 0.84 | 100% | 748 | 1 | $0.018 | 9 serena | 0.85 | 98% | 1368 | 5 | $0.065 | 29 codegraph | 0.93 | 100% | 1114 | 2 | $0.036 | 16 Now the tools **separate.** **grep collapses.** Lowest accuracy (0.71), only 83% of runs finish (the rest hit the 50-turn cap), and the ones that finish cost **6–24× more** ($0.42 vs $0.018–0.065) and take **6–18× longer** (~165s vs 9–29s). Text search drowns in noise when the question is "every call to this" or "which of a dozen identically-named methods." And the key bit: **graphlens — the mid-pack tool on easy tasks — is here the cheapest ($0.018) and fastest (9s).** Its semantic graph finally pays off: one call instead of twenty-seven. The most _accurate_ tool is codegraph (0.93). serena is competitive (0.85). So the same graphlens that looked unremarkable on pinpoint lookups becomes the most economical the moment the work is real — blast radius, refactoring. The ranking **inverts** between regimes. > Fairness note. MCP _resources_ are disabled for all arms. graphlens was the only server exposing resources, and in an early run the agent wandered into enumerating them and inflated cost ~24% until I denied them. All numbers above are from the clean re-run. ## Where the money goes: the mechanism is round-trips The cost difference is mostly **how many times the agent calls the tool** , which follows from how a server slices its primitives. On a simple "symbol → file" (`where_defined`), one call is enough for everyone. The gap opens on **relationship queries** — inheritance, route → handler, cross-language links. There `graphlens` chains fine-grained primitives (`find` → `neighbors` → `references`), while `codegraph` packs "source + call paths in one shot" (`explore` / `node`). This isn't a difference in _what the graph knows_ — graphs know roughly the same things. It's a difference in API granularity: fewer round-trips → cheaper and faster. That's why codegraph has the efficiency edge on simple tasks, and why grep bankrupts itself on hard ones — it makes 27 round-trips where the graph needs one or two. ## Model × tool interaction: the ranking drifts with model price This is the least obvious part. Take median $/task (across both regimes) broken down by model: Tool | haiku | sonnet | opus ---|---|---|--- filesystem | $0.053 | $0.080 | $0.087 graphlens | $0.020 | $0.041 | $0.046 serena | $0.026 | $0.033 | $0.042 codegraph | $0.023 | $0.041 | $0.031 Cheapest-first ranking **within each model** : * **haiku:** graphlens \$0.020 < codegraph \$0.023 < serena \$0.026 < filesystem \$0.053 * **sonnet:** serena \$0.033 < graphlens \$0.041 < codegraph \$0.041 < filesystem \$0.080 * **opus:** codegraph \$0.031 < serena \$0.042 < graphlens \$0.046 < filesystem \$0.087 Watch what happens to graphlens. On **haiku it's the cheapest of all.** On **opus it becomes the most expensive of the structural tools** (still cheaper than grep, though). The mechanism: graphlens results are **token-heavy** — graph neighborhoods, reference lists. On a cheap model that verbose context is nearly free; on an expensive one, opus prices the same tokens far higher, and verbosity hits the wallet. **serena and codegraph stay cheap on any model** because they return pinpoint results — they're robust to model choice; graphlens isn't. Which gives the most valuable takeaway of the lot: **a cheap model on a structural tool beats an expensive model on grep.** codegraph + haiku (~$0.023, accuracy ~0.99) beats filesystem + opus (~$0.087, accuracy 0.93) on every axis at once. ## The hypothesis that didn't hold I planted the two `xlang_link` tasks as a stress test: a TS call resolves to a Python handler across the `/api/v1/...` boundary, and I was sure single-language tools would trip on it. **They didn't.** Every arm, grep included, solved both cross-language tasks. The agent steps across the boundary itself, regardless of the context provider. On this set the hypothesis failed, and I report that as loudly as the findings that held. A benchmark that only reports what it hoped to see isn't a benchmark. ## Statistics, honestly Friedman test across the four tools, over task blocks, within each regime (df=3; critical values: 0.05 → 7.82, 0.01 → 11.34): SIMPLE: accuracy n=20 χ²= 0.40 (n.s.) — tie cost n=20 χ²=18.42 (p<.01) — serena < codegraph < graphlens < filesystem HARD: accuracy n= 6 χ²= 3.50 (n.s.) — underpowered cost n= 6 χ²=11.80 (p<.01) — graphlens < codegraph < serena < filesystem What's honest to claim from this: * The **cost difference is significant in both regimes** (p<.01). On HARD, graphlens is reliably the cheapest and grep reliably the most expensive. That's a solid result. * The **accuracy gap on HARD is large but not statistically significant** at n=6 (χ²=3.50). It's a strong _descriptive_ signal, not a proven one. Six tasks is few. * To firm up the accuracy claim you'd **add hard tasks, not cut simple ones.** Trimming the simple regime gives the hard one zero extra power — it just throws away good data. I'm leaving this in the article on purpose. The temptation to write "graphlens/codegraph are more accurate than grep, proven" is real, but n=6 doesn't carry it, and pretending otherwise would be dishonest. ## Index amortization: different currencies The structural tools build an index once — **pure static work, zero LLM tokens** , wall-clock only: Tool | one-time index ---|--- filesystem | 0s codegraph | 48s graphlens | 84s serena | 94s grep pays nothing up front but pays more per query. These are **different currencies** (seconds vs $/tokens), so I draw no single "break-even point" — that'd be a stretch. The picture is simple: the index is a one-time time cost with not a single token spent, while the $/task savings drip on every task. Over a long session the structural tools amortize; on a couple of one-off queries, grep's zero setup can win on time-to-first-answer. ## Takeaways for the business case Back to the original question: what do you feed the agent on a large project? **There is no "this tool is always best" answer.** There's a "depends on what work you hand it" answer: * **One-off pinpoint lookups** ("where is this class defined," "what does it inherit from"): use anything. grep keeps up, accuracy is the same, zero setup. You pay at most a small token overhead. * **Sustained blast-radius work** — refactoring, impact analysis, disambiguation on a large base: structural tools cut cost **6–24×** and latency **6–18×** vs grep — and, just as important, **they don't hit the turn cap.** grep on these tasks isn't just expensive; 17% of the time it never reaches an answer at all. * **Model choice interacts with tool choice.** A verbose graph is cheap on a small model and pricey on a big one. Running opus? Pick a tool with pinpoint output (codegraph, serena). Running haiku? graphlens is suddenly the cheapest. * **The cheapest combo isn't "expensive model + simple tool" — it's "cheap model + structural tool."** And the honest caveats, without which you can't transfer the conclusions to your project: > One repository (`apache/superset` @ 6.0.0), one harness, 26 tasks (20 simple / 6 hard). Regimes are reported separately and **never blended**. `cost_usd` is an API-equivalent, not a subscription bill. Failure = accuracy 0. This is **not a universal ranking** — it's a reproducible measurement on one concrete case. ## Where graphlens fits Since this is a follow-up to the graphlens post, let me say it straight. This benchmark does **not** prove graphlens is "the best." It shows the **specific regime where its structural graph pays off** (impact analysis, cheap and fast on cheaper models), and just as plainly shows **where it lags** (on opus its verbose output costs more than codegraph and serena; codegraph is more accurate on hard tasks). For me that's more useful than any victory lap. graphlens was built as an **engine and a precise polyglot graph model** , not a turnkey app — and the benchmark confirms exactly that: on structural questions the graph beats text search by a wide margin, and there's clear room to grow — MCP tool granularity (fewer round-trips, like codegraph) and output compactness (so it doesn't bankrupt itself on expensive models). That's my next work item, now backed by numbers instead of intuition. ## Reproduce it The whole harness and the raw data are open. A run reassembles deterministically from `data/`. * **Benchmark repo:** https://github.com/Neko1313/agent-context-bench * See `metrics.ipynb` (all charts and per-section stats) and `README.md` (methodology). * `uv run main.py` runs the full pipeline (clone superset → build indices → 936 runs, resumable within subscription limits), then open `metrics.ipynb`. If you've got a large project of your own and the itch to run the harness on it — issues and results welcome. The more independent runs across different codebases, the closer we get to an answer that transfers, rather than "works on superset." _(More on the tool itself: graphlens · docs.)_
dev.to
June 24, 2026 at 11:51 PM
graphlens: a polyglot code-analysis framework that turns your repo into a typed graph
# graphlens: turn any repo into one typed graph — across Python, TypeScript, Go and Rust Every code-intelligence tool I've ever used falls into one of two traps. The first is the **grep-and-read loop** : you (or your AI agent) search for a name, open ten files, read around the matches, follow an import, search again. It works, but it's slow, it burns tokens, and it has no idea that the `process_order` you found in `services.py` is the _same_ `process_order` that gets called from `api.py` — versus the unrelated one in `tests/`. The second is the **single-language silo** : tools that understand Python beautifully but go blind the moment your TypeScript front end calls a Python FastAPI route. Real systems are polyglot. Your tooling usually isn't. graphlens is an open-source (MIT) framework built to escape both traps. It parses a source project, normalizes its structure into a shared **graph IR** , and hands you that graph to do whatever you want with — dependency analysis, navigation, dead-code detection, or feeding an LLM agent precise answers instead of file dumps. Repository → Language Adapter → GraphLens (IR) → Graph Backend Layer | Responsibility ---|--- **Language Adapter** | Parses source files, produces a `GraphLens` **GraphLens** | Typed nodes + directed relations — the intermediate representation **Graph Backend** | Persists or queries the graph (Neo4j, in-memory, your own) The key design decision: **adapters are pure data producers.** They never write to a database, never touch the filesystem after reading, never run a server. The graph is the only output. That makes the whole pipeline trivially testable, cacheable, and serializable. ## 30 seconds to your first graph pip install "graphlens-cli[python]" graphlens analyze ./my-project graphlens · my-project nodes: 1240 relations: 3981 resolver: ok nodes by kind relations by kind FUNCTION 410 CONTAINS 980 METHOD 265 DECLARES 870 CLASS 98 CALLS 640 MODULE 54 REFERENCES 410 Or from Python: from pathlib import Path from graphlens import adapter_registry adapter = adapter_registry.load("python")() graph = adapter.analyze(Path("./my-project")) print(len(graph.nodes), "nodes,", len(graph.relations), "relations") fn = graph.nodes_by_name("process_order")[0] print("called by:", [n.name for n in graph.callers(fn.id)]) ## What makes the edges _real_ (and not name-matching guesses) Most lightweight code-graph tools resolve references by name: see a call to `save()`, draw an edge to anything called `save`. That's fast and wrong — there are usually a dozen `save`s in a codebase. graphlens splits the work in two: 1. **Tree-sitter** parses every file into a concrete syntax tree, giving exact structure and 1-based span positions. It records every _use-site_ as an **occurrence** with a role (call / read / write / annotation / base). 2. A language-specific, **type-aware resolver** then answers `definition_at(file, line, col)` for each occurrence. The resolved definition becomes a real edge to the _actual_ declaration node. Language | Resolver | Engine ---|---|--- Python | `TyResolver` | ty (Astral, Rust-based) via LSP TypeScript | `TsResolver` | the TypeScript Compiler API (Node subprocess) Go | `GoplsResolver` | gopls Rust | `RustAnalyzerResolver` | rust-analyzer So a `CALLS` edge points at the real function, a `HAS_TYPE` edge at the real class, an `INHERITS_FROM` edge at the real base. This is the difference between "probably related" and "is related". ### Honesty about partial failures Type analysis can degrade — a toolchain is missing, a file doesn't type-check. Instead of silently producing a half-resolved graph, graphlens records the outcome: from graphlens import RESOLVER_STATUS_KEY graph.metadata[RESOLVER_STATUS_KEY] # 'ok' | 'degraded' | 'unavailable' In CI you flip on `--strict` and a non-`ok` status fails the build, so an agent or dashboard never consumes a graph that's quietly incomplete. ## The graph model **Nodes** (`PROJECT`, `MODULE`, `FILE`, `CLASS`, `METHOD`, `FUNCTION`, `PARAMETER`, `VARIABLE`, `ATTRIBUTE`, `TYPE_ALIAS`, `IMPORT`, `DEPENDENCY`, `EXTERNAL_SYMBOL`, `BOUNDARY`) are frozen dataclasses with an id, kind, qualified name, file path, span, and free-form metadata. **Relations** are directed, typed edges: Kind | Meaning ---|--- `CONTAINS` / `DECLARES` | structural containment & declaration `IMPORTS` / `RESOLVES_TO` | import statements and where they resolve `CALLS` / `REFERENCES` / `INHERITS_FROM` / `HAS_TYPE` | resolved, type-aware edges `DEPENDS_ON` | declared package dependency `EXPOSES` / `CONSUMES` / `COMMUNICATES_WITH` | cross-language boundaries ### Deterministic IDs A node's ID is a SHA-256 hash of `project::kind::qualified_name`: from graphlens import make_node_id make_node_id("my-project", "my.module.func", "FUNCTION") # → the same id every scan, on every machine Because the ID depends only on identity, not file position, re-scanning yields the same IDs. That's what makes `graph.diff(other)` and incremental updates work — and what makes a graph cacheable in CI. ## The feature single-language tools can't have: cross-language boundaries This is my favorite part. Adapters emit language-agnostic **`BOUNDARY`** nodes for the interfaces a service exposes or consumes — HTTP routes, queue topics, gRPC methods, Temporal activities — with an `EXPOSES` edge (provider) or `CONSUMES` edge (consumer). A boundary's ID is `make_boundary_id(mechanism, key)` — _no project or language in it_. HTTP paths are normalized so that `/users/1`, `/users/{user_id}` (FastAPI), `<int:id>` (Flask), and `:id` (Express) all collapse to `GET /users/{}`. The payoff: a Python FastAPI route and a TypeScript `fetch` to the same endpoint produce the **same** boundary ID. Merge the two graphs, run `graphlens-link`, and you get `COMMUNICATES_WITH` edges spanning the language gap: from graphlens import adapter_registry from graphlens_link import link_graph py = adapter_registry.load("python")().analyze(python_project) ts = adapter_registry.load("typescript")().analyze(typescript_project) merged = py merged.merge(ts, allow_shared=True) # identical BOUNDARY nodes coincide result = link_graph(merged) # adds consumer → provider edges print(result.relations_added, "COMMUNICATES_WITH edges added") Now you can answer "which front-end calls hit this endpoint?" — a question no single-language tool can even represent. ## Five ways to use it **As a library** — load an adapter, get a `GraphLens`, query it: callers, callees, references, neighborhoods, diffs, JSON round-trips, multi-language merges. **From the CLI** — five subcommands cover the common workflows: graphlens analyze ./repo --output graph.json # index graphlens query process_order -g graph.json --op callers graphlens visualize ./repo # interactive vis.js HTML graphlens neo4j ./repo --uri bolt://localhost:7687 graphlens mcp --graph graph.json # serve to agents **In CI** — `--strict` plus a Docker image (`ghcr.io/neko1313/graphlens`) with every adapter and toolchain pre-installed. Index on every push, publish the graph as an artifact, fail on a degraded graph. **To LLM agents over MCP** — `graphlens mcp` exposes a saved graph as Model Context Protocol query tools (`stats`, `find`, `callers`, `callees`, `references`, `neighbors`, `boundaries`, `communicates_with`). Instead of dumping a codebase into the prompt, the agent asks precise questions and gets small structured answers — resolved edges, not best-effort text search. **As a Neo4j export** — straight into a graph database with `UNWIND … MERGE` Cypher (no APOC required), then query it however you like. ## Plugin architecture: the SQLAlchemy-dialect pattern The core never imports an adapter. Each language is a separate package that registers itself via Python entry points: [project.entry-points."graphlens.adapters"] python = "graphlens_python:PythonAdapter" Callers resolve adapters through a registry, by name string: adapter_registry.available() # ['python', 'typescript', ...] adapter = adapter_registry.load("python")() Adding a new language means writing one package against the `LanguageAdapter` contract — no changes to the core. ## What graphlens is _not_ The scope is deliberately narrow, and the docs spell it out. graphlens produces a graph IR and stops there. It does **not** : * persist state or own a database (backends are a separate consuming layer); * watch the filesystem or re-index incrementally on its own (scans are pure functions; deterministic IDs _enable_ incremental updates, but the caller drives them); * compute embeddings, semantic search, or relevance ranking (the graph is structural and type-aware, not a vector index); * provide a UI or an agent runtime (`visualize` emits static HTML, `mcp` exposes query tools — neither hosts a long-running service). Those belong to tools built _on top of_ graphlens. Keeping the core minimal is what keeps it composable. ## Benchmarks Throughput on real-world projects, refreshed on every release inside the published Docker image (single cold run, indicative): Project | Lang | LOC | Nodes | Time | Resolved ---|---|---|---|---|--- apache/superset | python | 399 519 | 156 251 | 148.7s | 84% colinhacks/zod | typescript | 74 194 | 8 741 | 19.0s | 91% gin-gonic/gin | go | 23 672 | 7 227 | 13.9s | 100% gohugoio/hugo | go | 224 821 | 34 809 | 112.7s | 99% BurntSushi/ripgrep | rust | 50 275 | 9 612 | 113.1s | 99% ## Try it pip install "graphlens-cli[python]" graphlens analyze . --output graph.json graphlens visualize . * **Repo:** https://github.com/Neko1313/graphlens * **Docs:** https://Neko1313.github.io/graphlens/ * **Requirements:** Python 3.13+. Python (`ty`) and TypeScript (Node) toolchains install on demand; Go and Rust adapters come via the Docker image. If you've ever wanted a single, accurate, language-agnostic model of "how does this codebase actually fit together" — that's exactly what graphlens hands you. I'd love feedback, issues, and adapter contributions.
dev.to
June 22, 2026 at 1:42 AM