Benchmarks local LLM inference speed (tokens/sec) on your own hardware via MCP tools.
Run one of the commands below, then add the client config underneath.
Pythonuvx inferbench-cli
Paste into Claude Desktop, Cursor (mcp.json), VS Code or any MCP client, then restart the client.
mcpServers{
"mcpServers": {
"inferbench": {
"command": "uvx",
"args": [
"inferbench-cli"
]
}
}
}
Inferbench is listed in the AI & ML category of the MCPNav directory. It is distributed as Python and can be loaded by any client that speaks the Model Context Protocol.
Typical uses include giving your assistant scoped access to the corresponding service so it can answer questions and take actions with real data instead of guessing. Always review what a server can access before you enable it — see our MCP security guide.
Inferbench is an MCP server by RudrenduPaul. Benchmarks local LLM inference speed (tokens/sec) on your own hardware via MCP tools.
Install it with: uvx inferbench-cli. Then add the JSON config to your client's MCP settings and restart the client.
The MCP server itself is free to install. The source is public on https://github.com/RudrenduPaul/InferBench. Any third-party API it calls (such as a search or maps API) may require its own key and billing.
Any MCP-compatible client can use it, including Claude Desktop, Cursor, VS Code, Windsurf and custom agents.