Governed GPU inference ops (vLLM + Ray Serve): latency RCA, scaling, drain, 39 tools.
Install
Run one of the commands below, then add the client config underneath.
uvx inference-aiopsClient configuration
Paste into Claude Desktop, Cursor (mcp.json), VS Code or any MCP client, then restart the client.
{
"mcpServers": {
"inference-aiops": {
"command": "uvx",
"args": [
"inference-aiops"
]
}
}
}About this server
Inference AIops is listed in the AI & Machine Learning category of the MCPNav directory. It is distributed as Python and can be loaded by any client that speaks the Model Context Protocol.
Typical uses include giving your assistant scoped access to the corresponding service so it can answer questions and take actions with real data instead of guessing. Always review what a server can access before you enable it — see our MCP security guide.
Frequently asked questions
What is the Inference AIops MCP server?+
Inference AIops is an MCP server by AIops-tools. Governed GPU inference ops (vLLM + Ray Serve): latency RCA, scaling, drain, 39 tools.
How do I install Inference AIops?+
Install it with: uvx inference-aiops. Then add the JSON config to your client's MCP settings and restart the client.
Is Inference AIops free to use?+
The MCP server itself is free to install. The source is public on https://github.com/AIops-tools/Inference-AIops. Any third-party API it calls (such as a search or maps API) may require its own key and billing.
Which clients support Inference AIops?+
Any MCP-compatible client can use it, including Claude Desktop, Cursor, VS Code, Windsurf and custom agents.