1. 为什么要在本地手写一个 MCP 客户端
MCP 全称 Model Context Protocol,是一个开源协议,用来标准化 LLM 与外部数据源、工具之间的交互方式。你可以把它理解成 AI 应用世界的 USB-C 接口:不管对面是 DeepSeek、Claude 还是别的模型,只要双方都按 MCP 说话,工具就能即插即用。它解决的问题很具体——过去每接一个外部能力(查时间、算 BMI、读数据库、调地图),都要为每个模型写一套适配代码;有了 MCP,工具方只写一次 Server,模型方只写一次 Client,中间靠协议对齐。
这篇面向的是想快速跑通 MCP 服务的开发者,尤其是习惯用 DeepSeek 这类 LLM 做工具调用的人。网上大多数 MCP 教程停留在 Claude Desktop、Cursor、Cline 这些现成客户端里点几下,但真实开发中你往往需要在自己的代码里构建客户端,把工具调用嵌进业务逻辑。所以本文从零开始:用 uv 初始化项目、写第一个 MCP Server、配置 TaoToken 统一 Key 与 API 通道、再写一个能真正发起工具调用的客户端,最后验证一次请求确实经 TaoToken 正常返回。
适合谁:会一点 Python、听说过 function calling 但没动手写过 MCP 的人;手里有多个模型 Key、想统一管理入口的人;想把本地函数暴露给 LLM 调用的人。全程大约 10 分钟,命令和配置都可直接复制。我试过把同一套 Server 分别接到不同模型上,只要客户端里的 Base URL 和 Key 换一下就能跑,这也是后面要引入 TaoToken 统一 Key 的原因——省去每个模型单独配 Key 的麻烦。
先明确一个概念边界:MCP Server 负责“提供能力”,MCP Client 负责“连接模型与 Server”。Server 里用@mcp.tool()装饰的函数会被自动解析成工具描述,包括函数名、注释、参数类型、返回类型,这些信息会拼进系统提示交给 LLM。LLM 判断该调哪个工具、填什么参数,返回一段 JSON,客户端解析后真正执行函数,再把结果回传给 LLM 生成自然语言回答。整条链路里,模型只负责“决策”,执行发生在你本地,安全边界清晰。
2. 用 uv 初始化项目并接入 TaoToken 统一 Key
环境管理这块,MCP 官方推荐 uv。uv 是用 Rust 写的高性能 Python 包安装器和虚拟环境管理器,目标是统一替代 pip、pip-tools、venv、virtualenv。它主要维护两个文件:pyproject.toml定义项目依赖和元信息,uv.lock记录完整依赖树且跨平台一致,后者由 uv 自动管理,别手动改。安装 uv 一条命令:
pip install uv然后创建项目并进入目录:
uv init mcp-server-demo cd mcp-server-demo把 MCP 的 Python SDK 加进依赖:
uv add "mcp[cli]"如果你用的是纯 pip 项目,等价写法是pip install "mcp[cli]"。日常几个高频命令记一下:uv add <包名>添加依赖并更新配置,uv pip install <包名>只装不写配置,uv run 文件名.py在项目环境里跑代码。后面所有脚本都用uv run启动,避免手动激活虚拟环境。
接下来是统一 Key 的部分。开发时最烦的是每个模型一个 Key、一个 Base URL,散落在各处。TaoToken 提供统一入口,把模型通道收敛成一套配置。你需要在控制台创建一个 API Key,然后所有请求都走同一个 Base URL。控制台地址是 https://taotoken.net/console ,创建 Key 的页面在 https://taotoken.net/api-keys 。拿到 Key 后,建议放进.env文件而不是硬编码:
# .env TAOTOKEN_API_KEY=sk-你的key TAOTOKEN_BASE_URL=https://taotoken.net/api客户端里用AsyncOpenAI指向这个 Base URL 即可,模型名按你实际要用的填,比如deepseek-chat。这样 Server 端完全不用关心模型是谁,Client 端只认一个入口。如果你更习惯用现成的编码工具,TaoToken 也提供 Coding Plan 通道,适合长期跑 Agent 场景,入口在 https://taotoken.net/coding-plan 。模型对话调试可以在 https://taotoken.net/models 里先验证通道是否通,接入文档在 https://taotoken.net/doc 。
这里给一份可直接复制的客户端配置片段,路径和字段名保持原样,方便你对照修改:
{ "base_url": "https://taotoken.net/api", "api_key": "sk-替换成你的key", "model": "deepseek-chat", "timeout": 60 }注意:Base URL 只写到
/api,不要在后面拼/v1之类的路径,具体以接入文档为准。Key 不要提交到 Git,.env记得加进.gitignore。
3. 编写第一个 MCP Server 并暴露两个工具
Server 是提供能力的地方。MCP 能提供的东西有四类:资源 Resources(文件内容、数据库记录、图像等)、提示 Prompt(可复用的提示模板和工作流)、工具 Tools(LLM 可直接调用的函数)、采样 Sampling(让 Server 反向请求 LLM 生成结果)。入门阶段先聚焦 Tools,因为它最直观。
在项目根目录新建server.py:
from mcp.server.fastmcp import FastMCP import datetime mcp = FastMCP() @mcp.tool() def get_time() -> str: """获取当前系统时间""" return str(datetime.datetime.now()) @mcp.tool() def calculate_bmi(weight_kg: float, height_m: float) -> float: """根据体重(kg)和身高(m)计算BMI""" return weight_kg / (height_m ** 2) if __name__ == "__main__": mcp.run(transport='stdio')写工具时有三个细节决定 LLM 能不能正确调用。第一,函数注释必须写清楚,MCP 会自动解析注释作为工具描述,注释含糊模型就选错工具。第二,参数类型要标全,weight_kg: float这种写法会被解析成参数描述,模型输出的字符串也会被自动转成对应类型。第三,返回值类型也标上,-> float让协议知道返回结构。这三点做到位,工具描述就完整了。
transport='stdio'表示用标准输入输出通信,适合本地开发,客户端通过子进程方式拉起 Server。跑起来后它不会打印什么,安静等待客户端连接,这是正常的。你可以先用uv run server.py确认没有语法错误,能启动就说明 Server 侧没问题。
4. 写客户端:连接 Server、解析工具、发起一次真实调用
客户端是用户与 LLM 交互的地方,流程分五步:连接 Server、列出可用工具、把工具描述拼进系统提示、用户输入后让 LLM 决策、若返回工具调用则执行并把结果回传。新建client.py:
import asyncio import sys import json from typing import Optional from contextlib import AsyncExitStack from mcp import ClientSession, StdioServerParameters from mcp.client.stdio import stdio_client from dotenv import load_dotenv from openai import AsyncOpenAI load_dotenv() def format_tools_for_llm(tool) -> str: args_desc = [] if "properties" in tool.inputSchema: for param_name, param_info in tool.inputSchema["properties"].items(): arg_desc = f"- {param_name}: {param_info.get('description', 'No description')}" if param_name in tool.inputSchema.get("required", []): arg_desc += " (required)" args_desc.append(arg_desc) return f"Tool: {tool.name}\nDescription: {tool.description}\nArguments:\n{chr(10).join(args_desc)}" class MCPClient: def __init__(self): self.session: Optional[ClientSession] = None self.exit_stack = AsyncExitStack() self.client = AsyncOpenAI( base_url="https://taotoken.net/api", api_key="sk-替换成你的key", ) self.model = "deepseek-chat" self.messages = [] async def connect_to_server(self, server_script_path: str): server_params = StdioServerParameters( command="python", args=[server_script_path], env=None ) self.stdio, self.write = await self.exit_stack.enter_async_context( stdio_client(server_params)) self.session = await self.exit_stack.enter_async_context( ClientSession(self.stdio, self.write)) await self.session.initialize() response = await self.session.list_tools() tools = response.tools print("\n服务器中可用的工具:", [tool.name for tool in tools]) tools_description = "\n".join([format_tools_for_llm(tool) for tool in tools]) system_prompt = ( "You are a helpful assistant with access to these tools:\n\n" f"{tools_description}\n" "Choose the appropriate tool based on the user's question. " "If no tool is needed, reply directly.\n\n" "IMPORTANT: When you need to use a tool, you must ONLY respond with " "the exact JSON object format below, nothing else:\n" "{\n" ' "tool": "tool-name",\n' ' "arguments": {\n' ' "argument-name": "value"\n' " }\n" "}\n\n" "After receiving a tool's response:\n" "1. Transform the raw data into a natural, conversational response\n" "2. Keep responses concise but informative\n" "3. Focus on the most relevant information\n" "4. Use appropriate context from the user's question\n" "5. Avoid simply repeating the raw data\n\n" "Please use only the tools that are explicitly defined above." ) self.messages.append({"role": "system", "content": system_prompt}) async def chat(self, prompt, role="user"): self.messages.append({"role": role, "content": prompt}) response = await self.client.chat.completions.create( model=self.model, messages=self.messages, ) return response.choices[0].message.content async def execute_tool(self, llm_response: str): try: tool_call = json.loads( llm_response.replace("```json\n", "").replace("```", "")) if "tool" in tool_call and "arguments" in tool_call: response = await self.session.list_tools() tools = response.tools if any(tool.name == tool_call["tool"] for tool in tools): try: print("[提示]:正在执行函数") result = await self.session.call_tool( tool_call["tool"], tool_call["arguments"]) print(f"[执行结果]: {result}") return f"Tool execution result: {result}" except Exception as e: error_msg = f"Error executing tool: {str(e)}" print(error_msg) return error_msg return f"No server found with tool: {tool_call['tool']}" return llm_response except json.JSONDecodeError: return llm_response async def chat_loop(self): print("MCP 客户端启动") print("输入 /bye 退出") while True: prompt = input(">>> ").strip() if prompt.lower() == '/bye': break llm_response = await self.chat(prompt) print(llm_response) result = await self.execute_tool(llm_response) if result != llm_response: self.messages.append( {"role": "assistant", "content": llm_response}) final_response = await self.chat(result, "system") print(final_response) self.messages.append( {"role": "assistant", "content": final_response}) else: self.messages.append( {"role": "assistant", "content": llm_response}) async def main(): if len(sys.argv) < 2: print("Usage: uv run client.py <path_to_server_script>") sys.exit(1) client = MCPClient() await client.connect_to_server(sys.argv[1]) await client.chat_loop() if __name__ == "__main__": asyncio.run(main())启动命令:
uv run client.py ./server.py启动后你会看到“服务器中可用的工具: ['get_time', 'calculate_bmi']”,说明客户端成功连上 Server 并解析出工具。接着输入“现在几点了”,模型会返回一段 JSON:
{ "tool": "get_time", "arguments": {} }客户端识别到工具调用,打印“[提示]:正在执行函数”,执行后回传结果,模型再生成“现在的时间是……”这样的自然语言。再试“身高180,体重80”,模型会自动把厘米转成米,调用calculate_bmi,返回 BMI 约 24.69 并判断属于正常范围。整个过程中,模型请求全部经 TaoToken 的 Base URL 发出,你可以在控制台看到调用记录,确认通道正常。
5. 常见报错排查:401、local proxy failed、reading choices
跑不通时先别怀疑代码,八成是配置或环境问题。下面按真实报错对照排查。
401 Unauthorized:最常见。原因通常是 Key 没填、填错、或者.env没被加载。检查client.py里api_key是否替换成了真实 Key,.env是否在项目根目录且被load_dotenv()读到。如果 Key 是从控制台复制的,注意别带多余空格。还有一种情况是 Base URL 写错,比如多写了/v1,导致鉴权路径不匹配,也会返回 401。确认 Base URL 就是https://taotoken.net/api。
local proxy failed / connection error:这类报错一般是网络层没通。先确认本机能否正常访问外网,再确认 Base URL 拼写无误。如果你在公司网络里,可能有出口限制,换一个网络环境试试。另外timeout设太短也会表现为连接失败,把超时调到 60 秒以上。
reading choices 相关报错:典型信息是'NoneType' object has no attribute 'choices'或读取response.choices[0]时报错。这通常意味着 API 返回结构不是预期的 chat completion 格式,可能是模型名写错、通道不支持该模型、或者返回了错误对象。先打印完整response看结构,再核对模型名是否与通道支持的列表一致。模型列表可以在 https://taotoken.net/models 里查。
OAuth / 鉴权跳转类报错:如果你用的是某些需要 OAuth 的客户端工具,报错提示授权失败,通常是回调地址或 token 过期。这类场景建议改用 API Key 方式接入,避免 OAuth 流程。TaoToken 的 API Key 方式在 https://taotoken.net/api-keys 创建,接入文档在 https://taotoken.net/doc 有完整说明。
工具没被调用:模型直接回答而不返回 JSON。检查系统提示里工具描述是否完整,函数注释是否为空。注释为空时工具描述就是空的,模型无从判断。另外确认format_tools_for_llm正确解析了inputSchema,参数描述缺失也会让模型犹豫。
Server 启动即退出:uv run server.py一闪而过。确认mcp.run(transport='stdio')在__main__里,且没有其他阻塞代码。stdio 模式下 Server 靠标准输入等待,客户端没连上时它安静挂着是正常的,不是卡死。
排查顺序建议:先单独跑 Server 确认能启动,再跑 Client 看能否列出工具,最后发一次对话看模型是否返回 JSON。每一步的输出都打印出来,定位会快很多。
6. 把统一 Key 用起来:从 Demo 到可扩展的工具链
跑通上面这套之后,你已经有了一个最小可用的 MCP 工具链。接下来可以做的扩展方向很多:把get_time换成查数据库、把calculate_bmi换成调内部 API,Server 侧只改函数,Client 侧几乎不用动。这就是 MCP 的价值——工具和模型解耦。
统一 Key 的好处在这个阶段会越来越明显。当你同时接多个模型做对比、或者在不同项目里复用同一套工具时,不用再维护一堆 Key 和 Base URL。所有请求走同一个入口,调用记录集中,排查问题也方便。如果你要长期跑编码类 Agent,Coding Plan 通道在 https://taotoken.net/coding-plan 有更合适的配额方案;日常调试模型对话用 https://taotoken.net/models 就够。
最后留一个实用技巧:把.env里的 Key 和 Base URL 抽成配置类,Client 初始化时统一读取,这样换环境只改一处。Server 脚本路径也建议用绝对路径,避免uv run client.py ./server.py在不同目录下找不到文件。工具函数的注释尽量写成人能看懂的一句话,模型选工具的准确率会明显提升。