Install LlamaIndex’s MCP integration package, connect a BasicMCPClient to your server, convert the server’s tools with McpToolSpec (or aget_tools_from_mcp_url), and pass those tools to a FunctionAgent. If you need to publish your own LlamaIndex workflow, expose it with workflow_as_mcp instead. The result is ordinary LlamaIndex tools that an agent can select alongside any native tools.
This guide covers URL and Streamable HTTP connections, local servers, OAuth, tool allow-lists, hosted LlamaIndex endpoints, workflow publishing, operations, and failure recovery.
Choose whether LlamaIndex will consume or publish MCP
MCP (Model Context Protocol) defines a standard way for an AI client to discover and call tools. In LlamaIndex, there are two distinct jobs:
- Consume: LlamaIndex connects to an existing MCP server, discovers its tools, and gives the converted tools to a
FunctionAgent. - Publish: You turn a LlamaIndex
Workflowinto an MCP server that other MCP clients can call.
The transport and hosting choices are independent of the job. A server can run locally behind a command, on your network, or over HTTP (including Streamable HTTP). Authentication can be absent, token-based, or OAuth.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchPC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11#1 Best Overall
| Decision | Consume an MCP server | Publish a LlamaIndex workflow |
|---|---|---|
| Primary API | BasicMCPClient, McpToolSpec, or aget_tools_from_mcp_url |
workflow_as_mcp |
| Direction | External tools become LlamaIndex tools | Your workflow becomes an MCP app |
| Governance | Use allowed_tools to restrict exposure |
Define the workflow’s public inputs and server configuration |
| Authentication | URL credentials, tokens, or BasicMCPClient.with_oauth |
Configure authentication in the MCP hosting layer |
Install the supported package
Install the MCP tools package alongside LlamaIndex. Add the OpenAI integration if you use the model shown in the examples:
pip install llama-index llama-index-tools-mcp llama-index-llms-openai
Set your model provider’s API key in the environment rather than hard-coding it:
export OPENAI_API_KEY="your-openai-key"
The MCP-specific package is llama-index-tools-mcp. Its documented classes and helpers are imported from llama_index.tools.mcp.
Connect to an MCP server with McpToolSpec
McpToolSpec is the flexible form: create a client, fetch the server’s tool definitions asynchronously, then pass the resulting list to a LlamaIndex agent.
import asyncio
import os
from llama_index.core.agent import FunctionAgent
from llama_index.llms.openai import OpenAI
from llama_index.tools.mcp import BasicMCPClient, McpToolSpec
async def main() -> None:
client = BasicMCPClient("https://example.com/mcp")
tool_spec = McpToolSpec(client=client)
tools = await tool_spec.to_tool_list_async()
agent = FunctionAgent(
llm=OpenAI(model="gpt-4.1", api_key=os.environ["OPENAI_API_KEY"]),
tools=tools,
system_prompt="You are an assistant with MCP tools. Use a tool when it is appropriate.",
)
response = await agent.achat("Use the MCP tools to answer my request.")
print(response)
if __name__ == "__main__":
asyncio.run(main())
Replace the example URL with the server’s MCP endpoint. The call to to_tool_list_async() performs discovery and conversion; after that, the agent sees the remote capabilities as normal LlamaIndex tools. Keep discovery inside your application startup path so an unavailable server fails visibly instead of leaving an agent with a silently incomplete tool set.
Limit discovery with allowed_tools
If a server offers more tools than this agent should use, allow-list names during discovery:
from llama_index.tools.mcp import aget_tools_from_mcp_url
tools = await aget_tools_from_mcp_url(
"http://127.0.0.1:8000/mcp",
allowed_tools=["search", "read_document"],
)
Only the named tools are returned. Use this as a least-privilege boundary: it reduces the model’s choices and prevents accidental access to administrative or destructive operations. The names must match the MCP server’s advertised tool names.
Rank #2
Use the direct URL helper
For a one-off connection, the helper is shorter than explicitly constructing a tool specification:
Free tools Windows power users keep installed
One-click scans. No signup required.
import asyncio
from llama_index.tools.mcp import aget_tools_from_mcp_url
async def load_tools():
return await aget_tools_from_mcp_url(
"https://example.com/mcp",
allowed_tools=["tool1", "tool2"],
)
tools = asyncio.run(load_tools())
Pass the returned tools list to FunctionAgent exactly as in the previous example. Use McpToolSpec when you need to retain the client object, configure OAuth, or make the discovery step part of a larger lifecycle.
Connect to local, HTTP, and Streamable HTTP servers
Local development
A local MCP server is useful while developing a tool. Start that server using its documented command, then point BasicMCPClient at the local HTTP endpoint, such as http://127.0.0.1:8000/mcp. Keep the server process running for the entire agent session; discovery can succeed while a later call fails if the process exits.
Hosted and Streamable HTTP
BasicMCPClient supports URL-based connections, including Streamable HTTP. Use the HTTPS endpoint supplied by the server operator, verify its certificate normally, and make sure outbound traffic from your runtime can reach that host. A reverse proxy must preserve the MCP route and the streaming response rather than buffering it indefinitely.
Headers and credentials
Do not put long-lived secrets in prompts or tool arguments. Configure the server’s expected authorization mechanism in the client or hosting environment, use short-lived credentials where possible, and redact authorization headers from logs. If your server requires a specific custom header, consult that server’s client configuration and keep the value in an environment variable or secret manager.
Recommended Free Tools
Add OAuth authentication
The MCP integration provides BasicMCPClient.with_oauth(...) for OAuth-protected servers. The documented setup accepts a client name, redirect URIs, a redirect handler, a callback handler, and optional token storage. When you omit custom token storage, the default is in memory.
from llama_index.tools.mcp import BasicMCPClient, McpToolSpec
client = BasicMCPClient.with_oauth(
"https://example.com/mcp",
client_name="my-llamaindex-agent",
redirect_uris=["http://127.0.0.1:8080/oauth/callback"],
redirect_handler=your_redirect_handler,
callback_handler=your_callback_handler,
# token_storage=your_token_storage, # optional persistent storage
)
tools = await McpToolSpec(client=client).to_tool_list_async()
Implement the handlers for your application’s browser or device flow. In production, supply durable, encrypted token storage if users should not authenticate again after every process restart; the in-memory default disappears when the process ends.
Rank #3
Use LlamaIndex’s hosted documentation MCP server
LlamaIndex publishes a documentation endpoint at https://developers.llamaindex.ai/mcp. It exposes search_docs, grep_docs, and read_doc. You can wrap it in the same specification used for any other server:
from llama_index.tools.mcp import BasicMCPClient, McpToolSpec
client = BasicMCPClient("https://developers.llamaindex.ai/mcp")
tools = await McpToolSpec(client=client).to_tool_list_async()
For a documentation assistant, provide a system prompt that tells the model to search and read the relevant documentation before answering, and optionally allow-list only those three documentation tools.
Run the official LlamaCloud MCP package
LlamaIndex also publishes the TypeScript package @llamaindex/llama-cloud-mcp. Its documented direct execution pattern is:
export LLAMA_CLOUD_API_KEY="your-llamacloud-key"
npx -y @llamaindex/llama-cloud-mcp
You can register that command in an MCP client such as Cursor, VS Code, or Claude Code. If your LlamaIndex agent must consume it over HTTP, run it behind an MCP-compatible service or use the connection mode supported by your deployment; the package’s client registration instructions determine the exact command and environment passed to that client.
Expose a LlamaIndex Workflow as an MCP server
Publishing is the reverse direction. Use workflow_as_mcp from llama_index.tools.mcp.utils to wrap a LlamaIndex Workflow. The utility accepts optional workflow_name, workflow_description, start_event_model, and additional FastMCP constructor arguments.
from llama_index.core.workflow import StartEvent, StopEvent, Workflow, step
from llama_index.tools.mcp.utils import workflow_as_mcp
class SummaryWorkflow(Workflow):
@step
async def summarize(self, ev: StartEvent) -> StopEvent:
text = getattr(ev, "text", "")
return StopEvent(result=text[:200])
workflow = SummaryWorkflow()
mcp_app = workflow_as_mcp(
workflow,
workflow_name="summary-workflow",
workflow_description="Return a short version of supplied text.",
)
Define the workflow’s start-event fields so the generated MCP interface has a clear input contract. Install the CLI extras when your chosen MCP serving path requires them:
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
pip install "mcp[cli]"
Then launch the returned MCP app using the serving command documented by your MCP runtime. Treat the resulting endpoint as a public API: validate inputs in workflow steps, authenticate callers, set timeouts, and avoid exposing internal state or unrestricted file and network access.
Operational design: transport, authentication, governance, and hosting
- Transport: choose a local command for development or HTTP/Streamable HTTP for a separately deployed service.
- Authentication: use no authentication only on a trusted local boundary; use tokens or OAuth for shared or internet-reachable servers.
- Tool governance: use
allowed_toolsfor client-side filtering, and keep dangerous operations off the server unless they are explicitly required. - Hosting: self-host when you control data locality and dependencies; use a LlamaIndex-hosted endpoint when its documentation or cloud capabilities match your use case.
Cache the discovered tool list for the lifetime of a process when the server’s tool schema is stable. Refresh it after deployments that add or remove tools. Set network and model timeouts independently so a slow remote tool does not consume an unlimited agent run.
Troubleshooting common failures
Package or import errors
Symptom: ModuleNotFoundError: llama_index.tools.mcp.
Cause: the MCP package is not installed in the active virtual environment.
Fix: run pip install llama-index-tools-mcp with that environment activated, then verify the interpreter used to launch the script is the same one where pip installed it.
Connection refused or 404
Symptom: discovery fails immediately.
Cause: the server is stopped, the path is wrong, or a proxy does not forward the MCP route.
Fix: confirm the process is running, copy the exact endpoint path (often ending in /mcp), and test reachability from the machine running LlamaIndex.
Tools are missing
Symptom: the agent cannot call an expected operation.
Cause: allowed_tools excluded it, or the server did not advertise it during discovery.
Fix: remove the allow-list temporarily to inspect the advertised names, then add the exact required names back.
OAuth callback never completes
Symptom: the browser returns to the wrong page or the client waits indefinitely.
Cause: the redirect URI registered with the provider does not exactly match the URI supplied to with_oauth, or the callback handler is not running.
Fix: register the precise scheme, host, port, and path; keep the callback listener reachable during authorization; persist tokens if the process must survive restarts.
Agent answers without using a tool
Symptom: the model responds from its own context.
Cause: the system prompt does not require tool use, the requested information does not need a tool, or the tool schema is unclear.
Fix: inspect the converted tool list, improve descriptions and input models on the server, and state when the agent must call a specific tool before answering.
Or skip the browser setup
If your workflow needs dependable website screenshots—for example, to give an agent visual evidence—ScreenshotNeo provides an HTTP API and MCP server. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; those cleanup steps can be disabled individually. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and each response identifies the result with X-Page-Verdict and X-Billed headers.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Scan for outdated or missing drivers - takes under a minute3Clear out junk files and repair common Windows errorsOne request is enough:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the complete parameter reference in the ScreenshotNeo documentation. The same endpoint supports PNG, JPEG, WebP, or PDF output and options such as full-page lazy-image loading, CSS-selector element capture, device and viewport presets, custom JavaScript and CSS, waits, request blocking, cookies and headers, geolocation, signed links, asynchronous webhooks, bulk capture of up to 100 URLs per call, and a usage API. An MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.
Best Value
ScreenshotNeo’s free plan includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account to get an API key.
FAQ
Can one LlamaIndex agent use native and MCP tools together?
Yes. Build one list containing your regular LlamaIndex tools and the tools returned by McpToolSpec, then pass that combined list to FunctionAgent. Keep names and descriptions distinct so the model can choose reliably.
Should tool discovery happen on every user request?
Usually no. Discover at startup or when the server schema changes, reuse the resulting tool definitions, and reconnect or refresh after a server deployment or connection failure.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Is OAuth token storage persistent by default?
No. BasicMCPClient.with_oauth uses in-memory token storage unless you provide a custom storage implementation, so a process restart normally requires another authorization.
Frequently Asked Questions
Can one LlamaIndex agent use native and MCP tools together?
Yes. Combine the native tools with the list returned by McpToolSpec before creating the FunctionAgent.
Should tool discovery happen on every user request?
Usually no. Discover at startup and refresh when the server’s tool schema changes or a connection fails.
Is OAuth token storage persistent by default?
No. The default storage is in memory; provide custom storage when tokens must survive restarts.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

