Quickstart¶
Install → connect → index → first grounded answer → see the savings. About five minutes.
1. Install¶
For containers, Windows Service, systemd, air-gap, see Deployment.
2. Connect your MCP client¶
Point your client at the sankshep command in stdio mode. Full per-client configs are on
Install; the VS Code shape:
// .vscode/mcp.json (or your user mcp.json)
{ "servers": { "sankshep": { "command": "sankshep", "args": ["serve", "--repo", "${workspaceFolder}"] } } }
--repo is the root every relative path resolves against. Reload the window after editing.
3. Confirm the tools are available — Agent mode, not Ask¶
MCP tools only fire when the client is in an agent/tool-using mode. In VS Code that's Copilot Chat's
Agent mode (not "Ask"). Ask the model to "list the sankshep tools" — you should see get_context,
search_code, index_repo, summarize_repo, remember, recall, export_decisions, token_report, and
the compose_task_prompt prompt. If not, see Troubleshooting.
4. Build the index (for semantic search)¶
get_context and summarize_repo work immediately — they read files directly. search_code needs an index:
index_repo (no arguments) → index_repo: indexed N file(s) under <repo>; the index now holds M chunk(s) in total.
The first index_repo downloads the local embedding model (~127 MiB, once) — expect a short one-time
delay; fully offline afterwards. When your client sends a progress token, index_repo reports one
notification per file, so a first index of a large repository is distinguishable from a hang.
A path that matches nothing, holds no supported source, or names a file rather than a directory returns an
error — not a silent success. So does search_code against an index that has never been built: it says
the index is empty and tells you to run index_repo, rather than returning an empty result you might read
as "that code does not exist".
Upgrading from 1.x?
Your existing index is discarded and rebuilt once, on first start. See Upgrading to 2.0.0, then Upgrading to 3.0.0.
5. Ask a grounded question¶
In Agent mode, let the model call get_context for you:
Use
get_contextonsrc/to explain how orders are validated, and cite the files.
Or force it explicitly with the tool. A good first call:
The result leads with a header — original -> delivered (% smaller), what it searched, what it withheld —
then the minimized code. The model answers from that.
6. See the savings¶
token_report → { "calls": …, "selectedTokens": …, "deliveredTokens": …, "compressionPct": …, "byTool": [ … ] }
compressionPct is delivered vs. the original size of the files delivered — a real, bounded number, no dollar
estimate. To measure savings and answer-quality together on a named repo, see Benchmarks.
Next¶
- Tool reference — every tool's arguments + real captured I/O.
- Prompt composer — turn a task + paths into a grounded prompt.
- Troubleshooting — when something doesn't behave.
- Security & privacy — what stays on your machine, and where.