Skip to main content

Overview

Beyond the function tools you define yourself, you can give a model new capabilities by connecting it to a remote Model Context Protocol (MCP) server. The model calls that server’s tools to reach and control external services when it needs them to answer a prompt. The mcp tool connects a user-supplied remote MCP server to an Agent API request. Agent API discovers the server’s tools when the request starts and calls them like native tools during the run, so you don’t have to write a custom function tool for each one. You can connect your own MCP server in two ways:
  • Use type: "mcp" to provide the server URL and authentication in each request. This requires Streamable HTTP.
  • Add a custom connector to save the server URL and authentication once in the API Console. Perplexity stores the MCP server’s token, so your application does not need to store or send it with each request. Any API key in that Project can then use it with type: "connector" and its connector ID. Custom connectors support Streamable HTTP and SSE.
Some managed connectors, such as GitHub, also provide credentials for Sandbox commands. Custom connectors and the mcp tool do not provide this Sandbox integration. The example below connects to the public DeepWiki MCP server, which needs no authentication, and asks the model to answer a question about a GitHub repository using the server’s tools.
For a fuller, runnable example that combines an MCP server with the model’s own web search, see the Model Picker cookbook recipe.
The sample below shows only the response’s output array, with long MCP tool outputs truncated.

Defer tool definitions

By default, every MCP tool definition the request exposes enters the model’s initial context. Set defer_loading to true when a server has many tools or large schemas, or when you connect several servers at once. The model can then search the catalog and load only the schemas it needs. The field is per server, so set it on each one you want deferred. Omitting the field, or setting it to false, keeps the default eager behavior. Deferred loading spends extra model turns before the first tool call. Set max_steps high enough that the model can search the catalog and still call the tools it finds. Otherwise a run can end right after the search and answer without ever calling a tool. In the example below, none of DeepWiki’s three tool definitions start in the model’s context. The model searches the catalog, loads only what the search matches, and calls the tool it found.
The sample below shows only the response’s output array, with long MCP tool outputs truncated.
If an MCP-heavy request approaches or exceeds the model’s context window, enable defer_loading before shortening the prompt or removing useful tools. This avoids placing every eligible MCP schema in the initial context while keeping those tools available to the model.
Each tool_search_output item is one search of the deferred catalog. The search can cover every deferred server or narrow to one, and it returns only the tools it matches — one of DeepWiki’s three above. A search can also return no match, and the model can then search again, so a run may hold several of these items. Calls still appear as mcp_call. Agent API still discovers each server’s tools when the request starts, so deferred loading does not eliminate discovery time or change discovery failure behavior. Three small tools is a modest catalog, so the extra search step buys little here. Deferred loading pays off as the catalog grows: more servers, more tools, larger schemas, or a context window you would otherwise exceed.

Authentication

Unlike the DeepWiki server above, most MCP servers require authentication. The most common scheme is an OAuth access token, which you pass in the authorization field of the mcp tool:
This example uses the GitHub MCP Server. Create a GitHub personal access token with access to the repositories you want the model to inspect, and export it as GITHUB_MCP_TOKEN.

Approvals

By default, Agent API runs MCP tool calls without approval. Set require_approval to review calls before they run. Approvals help you control what data the model sends to an MCP server and which actions it takes. A call that needs approval creates an mcp_approval_request item in the response output, and the tool does not run. Review the requested tool and its arguments, then send an mcp_approval_response item in a new request. Run the example below. It sends the first request, prints the response, asks you to type true or false, and then sends your decision. If the model answers without calling a tool, the response has no approval requests, so the example prints the answer and stops:

Get an approval request

The first request sets require_approval to "always", so every DeepWiki call needs a decision. Its output contains an approval request with the proposed tool and arguments:

Approve or deny the call

To continue, send previous_response_id and an mcp_approval_response for each approval request. Send the same tools again, unchanged: Agent API does not reuse tool definitions from the previous response. The approved call runs with the server_url, authorization, and headers from this request, not from the request that proposed the call. You do not need to send the mcp_approval_request or other earlier output items. The second request in the example has this body:
If you approve, Agent API runs the tool and returns an mcp_call whose approval_request_id matches the request. If you deny, the tool does not run, and the model answers without it. To tell the model why, add a reason string to a denial. Each approval applies to one tool call. A later response can contain new approval requests, so handle each one in the same way. To stop a pending call, deny it: removing the tool from allowed_tools does not cancel it.
Organizations with Zero Data Retention (ZDR) cannot use previous_response_id. Instead, send the conversation history in input: the original user message, the mcp_approval_request item, and your mcp_approval_response. An mcp_call item in input is history only and does not run the tool again. Send each mcp_call item back unchanged, including its approval_request_id: Agent API uses it to match the call to its approval request.

Choose which calls need approval

require_approval accepts these values: A filter can match tools by exact tool_names, by read_only, or by both. If you set both, a tool must match both. An empty filter {} matches every tool. If tool_names is omitted, null, or empty, the filter does not check the tool name. To require approval only for selected tools:
To run read-only tools without approval and require approval for all other tools:
The read_only value comes from the server’s tool annotations, shown in mcp_list_tools.tools[].annotations.read_only. The server sets this value, so it does not guarantee that a tool has no side effects.

Allow a tool for later calls

Apps often let the user choose between Allow once and Always allow. The API does not remember an approval, so store the Always allow choice in your application. Store the choice with the server’s server_url and the tool name. Do not rely on server_label alone: you set the label in each request, and the same label can point to a different server. In later requests, add the tool to the never filter:
Tools that are not in never still require approval. If the user has not chosen Always allow for any tool, do not send the never filter. An empty list, as in {"never": {"tool_names": []}}, matches every tool, so every call runs without approval. You can send the new policy in the same request as the mcp_approval_response. The pending call still needs its decision.

Handle approval errors

When you continue a response that has pending approval requests, the request fails with type: "invalid_request" if a decision is missing or invalid. No tool runs. The last error is a request validation error, so it fails with HTTP 400 in every mode. Agent API checks the other errors against the previous response when the run starts, so the request mode sets how they arrive: If the user moves on without deciding, deny the pending call. You can add the new user message after the mcp_approval_response in the same input.

Review calls safely

  • Before you approve, show the user the server_label, name, and parsed arguments.
  • Make the approval decision in your backend. Do not forward approval items from a browser or another untrusted client.

Parameters

Response shape

When an mcp tool is used, the response output array can include these MCP-specific item types alongside the final message item:
  • mcp_list_tools — emitted once per server, listing the tools discovered when the request starts.
  • mcp_call — emitted for each tool the model invokes on the server.
  • mcp_approval_request — during an approval flow, emitted instead of mcp_call when a proposed call needs a decision. The call has not run.
With defer_loading: true, the array can also include tool_search_output items when the model searches the deferred catalog. Tool invocations still appear as mcp_call items.

mcp_list_tools

Each entry in tools has the following fields:

mcp_call

mcp_approval_request

Emitted instead of mcp_call when a call needs approval. The call has not run.

tool_search_output

Emitted only with defer_loading: true, once per search the model runs against the deferred catalog. Each entry in tools is a namespace for one server: Example response output array:

Error handling

A discovery failure happens when a server cannot be reached or returns an unusable response as its tools are listed at the start of the run. Because discovery runs before the model, the whole run fails with external_connector_error. With stream: true or background: true, both failures arrive the same way: the HTTP status is 200, a stream ends with a response.failed event that holds the error, and a background response ends with status: "failed". Tool-call failures during the run do not fail the request. The error is returned to the model in-band on the mcp_call item (as above), so the model can recover or explain it in its final answer.

Risks and safety

The mcp tool lets you connect models to external services — a powerful capability that carries risk. Remote MCP servers are third-party services that have not been verified by Perplexity. They can let a model read, send, and receive data, and take actions in the connected service, and each server is subject to its own terms and conditions. Connect only servers you trust. Calls run without approval by default. For tools that write data or take actions, set require_approval so you can review each call before it runs. Use allowed_tools to limit which server tools the model can call. If the model only needs to read data, a read-only token or server mode is a simpler alternative.

Limitations

The mcp tool uses an OpenAI-compatible shape for many fields, but the APIs are not identical. These MCP features remain unsupported or behave differently:

Pricing

MCP tool calls are free — Agent API does not charge a per-invocation fee for calling a remote MCP server. Model token usage is still billed separately according to Agent API token pricing (see Models for per-model rates), and you operate the remote MCP server, so any cost it incurs is outside Agent API billing.