【免费下载链接】DeepSeek-V4-Pro-0813
导读:本文基于 DeepSeek-V4-Pro-0813 官方仓库 README 及其配套的
encoding与inference两个子模块,系统讲解该模型的核心能力、OpenAI 兼容消息的编码/解析方案(含工具调用、扩展思考与快速指令 token)、以及 vLLM、SGLang、本地多卡推理三种部署路径。读完本文,你将掌握如何把多轮对话与工具调用消息正确编码为模型输入、如何用单条启动命令开启 DSpark 投机解码加速,以及如何在本地完成权重转换与交互式聊天。
一、模型概览:从预览版到正式版的关键升级
DeepSeek-V4-Pro-0813是 DeepSeek-V4-Pro 的官方正式发布版本(仓库根目录 README.md),取代了此前的预览版本。它在生产环境中展现出显著增强的 Agent 能力与性能提升,其基础架构沿用 DeepSeek-V4-Pro (Preview) 的模型结构,并额外挂载了DSpark 投机解码(speculative decoding)模块。
根据仓库根目录 README.md 中的说明,该版本在仓库公布的多个基准上优于 DeepSeek-V4-Pro (Preview),并与最强的闭源模型整体处于同一水平。仓库官方公布的具体基准数据如下:
| Benchmark | DeepSeek-V4-Pro-0813 | DeepSeek-V4-Flash-0731 | DeepSeek-V4-Pro (Preview) | DeepSeek-V4-Flash (Preview) | GLM-5.2 | Kimi K3 | Opus-4.8 | Fable-5 (w/ fallback) |
|---|---|---|---|---|---|---|---|---|
| HLE (wo / w tools) | 42.7 / 60.0 | 37.8 / 51.5 | 37.7 / 48.2 | 34.8 / 45.1 | 40.5 / 54.7 | 43.5 / 56.0 | 49.8 / 57.9 | 53.3 / 63.0 |
| Terminal Bench 2.1 | 87.9 | 82.7 | 72.1 | 61.8 | 81.0 | 88.3 | 85.0 | 88.0 |
| NL2Repo | 61.5 | 54.2 | 38.5 | 39.4 | 48.9 | - | 69.7 | - |
| Cybergym | 83.3 | 76.7 | 52.7 | 38.7 | - | 80.0 | 78.3 | 83.1 |
| DeepSWE | 62.7 | 54.4 | 12.8 | 7.3 | 46.2 | 67.5 | 58.0 | 70.0 |
| Toolathlon-Verified | 74.1 | 70.3 | 55.9 | 49.7 | 59.9 | 76.5 | 76.2 | 77.9 |
| Agents' Last Exam | 25.7 | 25.2 | 16.5 | 15.8 | 23.8 | 27.6 | 25.7 | - |
| AutomationBench (Public) | 31.8 | 25.1 | 12.8 | 10.8 | 12.9 | 30.8 | 27.2 | 29.1 |
| DSBench-FullStack † | 71.1 | 68.7 | 41.8 | 37.0 | 61.8 | 73.7 | 71.6 | 77.2 |
| DSBench-Hard † | 67.2 | 59.6 | 31.1 | 25.8 | 54.5 | 63.0 | 71.7 | 68.3 |
评估说明(摘自仓库 README):
- 上述公开基准中的代码 Agent 任务,DeepSeek-V4-Pro-0813 使用 DeepSeek Harness 的最小模式(minimal mode)作为 Agent 框架评估,推理努力级别取
max,采样参数为temperature = 1.0, top_p = 0.95。 - † 标记的 DSBench-FullStack 为内部全栈开发测试集,DSBench-Hard 为内部困难编码 Agent 问题测试集。
1.1 模型架构关键参数
仓库根目录 config.json 给出了官方推理配置,几个关键参数可以帮你在部署前建立对模型规模的准确认知:
- 架构类型:
DeepseekV4ForCausalLM,model_type为deepseek_v4; - 上下文长度:
max_position_embeddings = 1048576(约 100 万 token),通过rope_scaling(yarn类型,factor = 16,原始长度 65536)扩展; - MoE 结构:
n_routed_experts = 384个路由专家、n_shared_experts = 1个共享专家、每 token 激活num_experts_per_tok = 6个专家,moe_intermediate_size = 3072,路由打分函数为sqrtsoftplus(scoring_func),topk_method = noaux_tc,routed_scaling_factor = 2.5; - 注意力:
num_attention_heads = 128、head_dim = 512、num_key_value_heads = 1(MQA 结构)、q_lora_rank = 1536、o_lora_rank = 1024、o_groups = 16、sliding_window = 128; - 量化:
quantization_config为动态 FP8(quant_method: fp8、activation_scheme: dynamic、weight_block_size: [128, 128]、fmt: e4m3、scale_fmt: ue8m0),专家权重默认expert_dtype: fp4; - DSpark 投机解码:
dspark_block_size = 5、dspark_target_layer_ids = [58, 59, 60]、dspark_markov_rank = 512、dspark_noise_token_id = 128799。
这些参数可以在 inference/model.py 的ModelArgs数据类中找到一一对应的字段,本地推理时会直接从config.json加载。
二、消息编码方案:OpenAI 兼容消息与模型输入字符串的转换
这是本版本最值得注意的工程细节:本次发布没有附带 Jinja 格式的 chat template。取而代之的是仓库提供了一个专门的encoding文件夹,内含 Python 脚本与测试用例,演示如何把 OpenAI 兼容格式的消息编码为模型的输入字符串,以及如何解析模型的文本输出。完整文档见 encoding/README.md。
官方给出的速览示例(摘自根目录 README.md):
from encoding_dsv4 import encode_messages, parse_message_from_completion_text messages = [ {"role": "user", "content": "hello"}, {"role": "assistant", "content": "Hello! I am DeepSeek.", "reasoning_content": "thinking..."}, {"role": "user", "content": "1+1=?"} ] # messages -> string prompt = encode_messages(messages, thinking_mode="thinking", reasoning_effort="max") # string -> tokens import transformers tokenizer = transformers.AutoTokenizer.from_pretrained("deepseek-ai/DeepSeek-V4-Pro-0813") tokens = tokenizer.encode(prompt)从 encoding/encoding_dsv4.py 的源码看,encode_messages是主入口,内部会依次完成:工具消息合并(merge_tool_messages)、工具结果排序(sort_tool_results_by_call_order)、BOS 插入、思考内容裁剪(_drop_thinking_messages)以及逐条消息渲染(render_message)。
2.1 特殊 Token 一览
| Token | 用途 |
|---|---|
<|begin▁of▁sentence|> | 序列开始(BOS) |
<|end▁of▁sentence|> | 助手回合结束(EOS) |
<|User|> | 用户回合前缀 |
<|Assistant|> | 助手回合前缀 |
<|latest_reminder|> | 最新提醒(日期、地区等) |
<think>/</think> | 推理块定界符 |
|DSML| | DSML 标记 token |
这些 token 在 encoding_dsv4.py 中作为模块级常量定义。
2.2 支持的消息角色
编码支持以下角色:system、user、assistant、tool、latest_reminder和developer。
需要注意两点:
developer角色仅用于内部搜索 Agent 流水线,通用聊天或工具调用任务不需要它,官方 API 也不接受该角色的消息;tool角色并非独立渲染——DeepSeek-V4 把工具结果合并进用户消息,以<tool_result>块形式呈现(源码中role == "tool"分支直接抛出NotImplementedError,并提示先用merge_tool_messages()预处理)。从 OpenAI 格式转换时,encode_messages会自动完成这一合并,无需手动处理。
2.3 基础多轮对话格式
<|begin▁of▁sentence|>{system_prompt} <|User|>{user_message}<|Assistant|></think>{response}<|end▁of▁sentence|> <|User|>{user_message_2}<|Assistant|></think>{response_2}<|end▁of▁sentence|>- BOS token 总是预置在对话最开头(
add_default_bos_token=True时); - 在chat 模式(
thinking_mode="chat")下,</think>紧跟在<|Assistant|>之后立即闭合思考块,模型直接生成正文内容。
2.4 交错思考模式(Thinking Mode)
在thinking 模式(thinking_mode="thinking")下,模型会在回答前于<think>...</think>块内产出显式推理过程:
<|begin▁of▁sentence|>{system_prompt} <|User|>{message}<|Assistant|><think>{reasoning}</think>{response}<|end▁of▁sentence|>drop_thinking参数(默认True)控制是否保留早期回合的推理内容,这是节省上下文的关键机制:
- 无工具场景:
drop_thinking生效。最后一个用户消息之前的助手回合推理内容会被剥离,只有最后一轮助手回合保留<think>...</think>块; - 有工具场景(system 或 developer 消息上定义了
tools):drop_thinking自动失效(encode_messages源码中检测到任一消息携带tools即把effective_drop_thinking置为False)。所有回合都保留推理内容,因为工具调用对话需要完整上下文,模型才能跨工具调用跟踪多步推理。
源码中_drop_thinking_messages的裁剪规则(encoding_dsv4.py):user、system、tool、latest_reminder角色始终保留;位于最后一个用户索引处及之后的助手消息保留推理;之前的助手消息移除reasoning_content;之前的developer消息整条丢弃。
2.5 工具调用(DSML 格式)
工具定义在system或developer消息的tools字段(OpenAI 兼容格式)上。当存在工具时,会向 system/user 提示注入如下 schema 块(该模板即源码中的TOOLS_TEMPLATE,encoding_dsv4.py):
## Tools You have access to a set of tools to help answer the user's question. You can invoke tools by writing a "|DSML|tool_calls" block like the following: |DSML|tool_calls |DSML|invoke name="$TOOL_NAME" |DSML|parameter name="$PARAMETER_NAME" string="true|false">$PARAMETER_VALUE|DSML|parameter ... |DSML|invoke |DSML|invoke name="$TOOL_NAME2" ... |DSML|invoke |DSML|tool_calls String parameters should be specified as is and set `string="true"`. For all other types (numbers, booleans, arrays, objects), pass the value in JSON format and set `string="false"`. If thinking_mode is enabled (triggered by <think>), you MUST output your complete reasoning inside <think>...</think> BEFORE any tool calls or final response. Otherwise, output directly after </think> with tool calls or final response. ### Available Tool Schemas {tool_definitions_json} You MUST strictly follow the above defined tool name and parameter schemas to invoke tool calls.一个真实的工具调用在助手回合中长这样:
|DSML|tool_calls |DSML|invoke name="function_name" |DSML|parameter name="param" string="true">string_value|DSML|parameter |DSML|parameter name="count" string="false">5|DSML|parameter |DSML|invoke |DSML|tool_calls<|end▁of▁sentence|>string="true":参数值是原始字符串;string="false":参数值是 JSON(数字、布尔、数组、对象)。
工具执行结果包裹在用户消息内的<tool_result>标签中:
<|User|><tool_result>{result_json}</tool_result><|Assistant|><think>...当存在多个工具结果时,它们会按照前一条助手消息中对应tool_calls的顺序排序(由sort_tool_results_by_call_order实现)。编码时,参数值按类型自动决定string标志:字符串参数置"true"原样输出,非字符串参数(数字、布尔、数组、对象)置"false"并以 JSON 序列化——这正是encode_arguments_to_dsml的逐参数处理逻辑(encoding_dsv4.py)。
2.6 快速指令特殊 Token
快速指令 token 用于辅助分类与生成类任务。它们通过消息的"task"字段追加到消息上,触发模型针对单 token 或短格式输出的专用行为(定义见 encoding_dsv4.py 的DS_TASK_SP_TOKENS):
| 特殊 Token | 描述 | 格式 |
|---|---|---|
<|action|> | 判断用户提示是否需要联网搜索或可直接回答。 | ...<|User|>{prompt}<|Assistant|><think><|action|> |
<|title|> | 在第一条助手回复后生成简洁对话标题。 | ...<|Assistant|>{response}<|end▁of▁sentence|><|title|> |
<|query|> | 为用户提示生成搜索查询词。 | ...<|User|>{prompt}<|query|> |
<|authority|> | 对用户提示的来源权威性需求进行分类。 | ...<|User|>{prompt}<|authority|> |
<|domain|> | 识别用户提示所属领域。 | ...<|User|>{prompt}<|domain|> |
<|extracted_url|><|read_url|> | 判断用户提示中的每个 URL 是否需要抓取阅读。 | ...<|User|>{prompt}<|extracted_url|>{url}<|read_url|> |
消息格式中的使用规则:
action位于用户消息上:<|action|>token 放在助手前缀与思考 token 之后,触发路由决策(例如 "Search" 或 "Answer");- 其他任务(
query、authority、domain、read_url)位于用户消息上:任务 token 直接追加在用户内容之后; title位于助手消息上:<|title|>token 追加在助手 EOS 之后,由下一条助手消息提供生成的标题。
2.7 推理努力级别(Reasoning Effort)
reasoning_effort参数支持low、high、max三个级别,控制模型在回答前投入的思考(deliberation)程度。该级别纯粹以文本前缀的形式实现——在 thinking 模式下,所选级别的前缀文本被预置到提示的最开头(system 消息之前),其余编码完全相同(encoding_dsv4.py 的REASONING_EFFORT_PROMPTS):
reasoning_effort | 提示前缀 |
|---|---|
"low"(默认) | 无 |
"high" | Reasoning Effort: Absolute maximum ... |
"max" | Reasoning Effort: Beyond maximum ... |
reasoning_effort在 chat 模式(thinking_mode="chat")下无效,因为该模式下模型根本不产出推理块。
"high"的完整前缀文本:
Reasoning Effort: Absolute maximum with no shortcuts permitted. You MUST be very thorough in your thinking and comprehensively decompose the problem to resolve the root cause, rigorously stress-testing your logic against all potential paths, edge cases, and adversarial scenarios. Explicitly write out your entire deliberation process, documenting every intermediate step, considered alternative, and rejected hypothesis to ensure absolutely no assumption is left unchecked."max"的完整前缀文本:
Reasoning Effort: Beyond maximum — exhaustive, relentless, and uncompromising. You MUST reason with the utmost depth and rigor, leaving absolutely nothing to chance: exhaustively decompose the problem into its most fundamental components, trace every causal chain to its root, and resolve the underlying cause rather than any surface symptom. Do not stop reasoning until you have independently verified the solution from multiple angles and are certain that no assumption remains unchecked and no error remains undiscovered.2.8 解析模型输出:parse_message_from_completion_text
parse_message_from_completion_text(text, thinking_mode)负责把模型单回合的原始输出文本解析为结构化助手消息,返回{"role": "assistant", "content", "reasoning_content", "tool_calls"}(tool_calls 为 OpenAI 格式)。解析流程(encoding_dsv4.py):thinking 模式下先读到</think>提取reasoning_content;随后读到 EOS 或 DSML tool_calls 起始标记提取content;若存在工具调用,则由parse_tool_calls逐个解析invoke/parameter块并还原为 JSON 参数。
重要限制:该函数仅面向格式良好的模型输出设计,不尝试纠正模型偶尔产生的畸形输出。生产环境中建议自行增加额外的错误处理。
三、编码方案的工程验证:测试用例与可复现实验
encoding/test_encoding_dsv4.py 提供了 4 个端到端测试用例,覆盖了编码 + 解析的完整闭环,测试输入输出存放在 encoding/tests 目录(test_input_N.json/test_output_N.txt)。直接运行即可验证:
python test_encoding_dsv4.py四个用例分别验证:
- thinking 模式 + 工具调用(多轮、工具结果合并进用户消息):解析出
get_weather工具调用,参数{"location": "Beijing", "unit": "celsius"},最终回合reasoning_content为 "Got the weather data...",正文包含 "22°C"; - thinking 模式无工具(
drop_thinking剥离早期推理):断言"The user said hello"不在最终 prompt 中,验证早期推理被正确丢弃; - 交错 thinking + 搜索(developer 携带工具、latest_reminder);
- 快速指令任务 + latest_reminder(chat 模式、action 任务)。
这些测试同时是理解编码行为的最快途径:例如用例 1 验证了多工具结果按调用顺序排序,用例 2 验证了drop_thinking的上下文压缩效果。
四、部署路径一:vLLM 推理服务与 DSpark 投机解码
DSpark 投机解码只需一个标志即可开启:在 vLLM 启动命令中追加--speculative-config,指定method: dspark:
--speculative-config '{"method":"dspark","num_speculative_tokens":7,"draft_sample_method":"greedy"}'官方在单个 4×GB300 节点上服务该模型的完整命令示例(摘自根目录 README.md):
vllm serve deepseek-ai/DeepSeek-V4-Pro-0813 \ --trust-remote-code --kv-cache-dtype fp8 --block-size 256 \ --data-parallel-size 4 --enable-expert-parallel \ --moe-backend deep_gemm_mega_moe \ --attention-config '{"use_fp4_indexer_cache": true}' \ --speculative-config '{"method":"dspark","num_speculative_tokens":7,"draft_sample_method":"greedy"}'各参数含义与配置要点:
--trust-remote-code:允许加载模型仓库中的远程代码(transformers 集成必需);--kv-cache-dtype fp8:KV 缓存使用 FP8 存储,显著降低大上下文下的显存占用;--block-size 256:PagedAttention 的 KV 块大小;--data-parallel-size 4与--enable-expert-parallel:配合 4 卡节点,开启数据并行与专家并行(EP);--moe-backend deep_gemm_mega_moe:使用 DeepGEMM 的 MoE 后端;--attention-config '{"use_fp4_indexer_cache": true}':启用 FP4 索引器缓存(对应配置中expert_dtype: fp4与 DSpark 索引机制);--speculative-config:开启 DSpark 投机解码,num_speculative_tokens = 7表示每步最多投机 7 个候选 token,draft_sample_method = "greedy"表示草稿采样采用贪心策略。
关于 DSpark 在模型侧的依据:仓库根目录 config.json 中的dspark_block_size = 5、dspark_target_layer_ids = [58, 59, 60]、dspark_markov_rank = 512等字段即 DSpark 模块的超参数,本地推理实现(inference/model.py 的ModelArgs)与 kernel(inference/kernel.py)中也包含对应的 DSpark 相关实现,目标层与主模型共享同一份 checkpoint,因此无需单独指定草稿模型路径。
五、部署路径二:SGLang 推理服务
SGLang 上启用 DSpark 的方式是使用--speculative-algorithm DSPARK,并且不要单独设置--speculative-draft-model-path——因为目标权重与草稿权重来自同一个 checkpoint(这正是 DSpark 设计的特点之一)。
官方示例命令(摘自根目录 README.md):
sglang serve \ --trust-remote-code \ --model-path deepseek-ai/DeepSeek-V4-Pro-0813 \ --tp 4 \ --moe-runner-backend flashinfer_mxfp4 \ --speculative-algorithm DSPARK \ --mem-fraction-static 0.90 \ --chunked-prefill-size 4096 \ --swa-full-tokens-ratio 0.1参数要点:
--tp 4:张量并行度 4(对应单节点 4 卡);--moe-runner-backend flashinfer_mxfp4:MoE 运行后端采用 FlashInfer 的 MXFP4 路径(与模型的 FP4 专家权重配合);--speculative-algorithm DSPARK:开启 DSpark 投机解码;--mem-fraction-static 0.90:静态显存占用比例上限;--chunked-prefill-size 4096:prefill 分块大小;--swa-full-tokens-ratio 0.1:SWA(sliding window attention)满 token 比例。
六、部署路径三:本地多卡推理(权重转换 + 交互式聊天)
本地运行的详细说明在 inference/README.md,核心流程分为两步:先转换权重,再启动推理。
6.1 第一步:转换 Hugging Face 权重
export EXPERTS=256 export MP=4 export CONFIG=config.json python convert.py --hf-ckpt-path ${HF_CKPT_PATH} --save-path ${SAVE_PATH} --n-experts ${EXPERTS} --model-parallel ${MP}说明:EXPERTS为模型总专家数(仓库配置中n_routed_experts = 384,此处以 README 示例的 256 为准,需与 checkpoint 一致),MP为模型并行度,HF_CKPT_PATH指向 Hugging Face 格式权重目录,SAVE_PATH为转换输出目录。转换脚本 inference/convert.py 会把*.safetensors按模型并行度切分,专家权重按idx分片到各 rank,并处理wo_a权重反量化合并;转换完成后还会把tokenizer.json、tokenizer_config.json一并复制到输出目录。
FP8 切换说明:若想使用 FP8 专家权重,只需在config.json中移除"expert_dtype": "fp4"字段,并在convert.py中追加--expert-dtype fp8。脚本的cast_e2m1fn_to_e4m3fn会把 FP4(e2m1fn)专家权重无损转换为 FP8(e4m3fn),并计算新的缩放因子。
6.2 第二步:启动推理
交互式聊天:
torchrun --nproc-per-node ${MP} generate.py --ckpt-path ${SAVE_PATH} --config ${CONFIG} --interactive从文件批量推理:
torchrun --nproc-per-node ${MP} generate.py --ckpt-path ${SAVE_PATH} --config ${CONFIG} --input-file ${FILE}多节点推理:
torchrun --nnodes ${NODES} --nproc-per-node $((MP / NODES)) --node-rank $RANK --master-addr $ADDR generate.py --ckpt-path ${SAVE_PATH} --config ${CONFIG} --input-file ${FILE}inference/generate.py 的实现要点:交互模式下会维护messages列表,每次用encode_messages(messages, thinking_mode="chat")编码后送入generate()做 prefill + decode,输出经parse_message_from_completion_text(completion, thinking_mode="chat")解析回结构化消息并追加到历史,实现多轮对话;支持/exit退出、/clear清空上下文;--max-new-tokens(默认 300)与--temperature(默认 1.0)可调。
6.3 本地部署的采样参数建议
对于本地部署,官方建议(摘自根目录 README.md):
- 采样参数设为
temperature = 1.0; - Agent 场景下
top_p = 0.95,其他场景top_p = 1.0; - 对于
high和max推理努力级别,建议最大输出长度设为384Ktoken。
仓库根目录 generation_config.json 的默认值(do_sample: true、temperature: 1.0、top_p: 1.0)与上述建议一致。
七、快速开始:编码 + 推理的最小可运行闭环
综合以上内容,一个最小可运行的端到端闭环如下:
from encoding_dsv4 import encode_messages, parse_message_from_completion_text # 1. 定义对话(OpenAI 兼容格式) messages = [ {"role": "system", "content": "You are a helpful assistant."}, {"role": "user", "content": "What is 2+2?"}, ] # 2. 编码为模型输入字符串 prompt = encode_messages(messages, thinking_mode="thinking") # => "<|begin▁of▁sentence|>You are a helpful assistant.<|User|>What is 2+2?<|Assistant|><think>" # 3. (在真实部署中此处调用推理引擎得到 completion) # 4. 解析模型输出为结构化消息 completion = "Simple arithmetic.</think>2 + 2 = 4.<|end▁of▁sentence|>" parsed = parse_message_from_completion_text(completion, thinking_mode="thinking") # => {"role": "assistant", "reasoning_content": "Simple arithmetic.", "content": "2 + 2 = 4.", "tool_calls": []}要点回顾:
- 编码入口:
encode_messages(messages, thinking_mode, reasoning_effort, drop_thinking, context, add_default_bos_token),其中thinking_mode取"chat"或"thinking",reasoning_effort取"low"(默认)/"high"/"max"; - 解析入口:
parse_message_from_completion_text(text, thinking_mode),仅适用于格式良好的输出; - 无 Jinja 模板:模型仓库不提供 Jinja chat template,任何生产集成都必须使用
encoding模块完成编解码,测试用例(encoding/tests)即官方行为规范。
八、许可与引用
本仓库及其模型权重采用MIT License(见 LICENSE)。若在学术工作中使用 DeepSeek-V4,可参考根目录 README 提供的引用格式:
@misc{deepseekai2026deepseekv4, title={DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence}, author={DeepSeek-AI}, year={2026}, }如有问题,可在仓库提交 issue 或联系 service@deepseek.com。
小结:DeepSeek-V4-Pro-0813 是一套面向 Agent 场景的正式发布模型。其部署链路的关键在于三件事:用encoding模块(而非 Jinja 模板)完成消息编解码、用--speculative-config/--speculative-algorithm DSPARK单参数开启 DSpark 投机解码、以及按官方建议的采样参数(temperature=1.0、Agent 场景top_p=0.95、high/max努力级别下 384K 最大输出)配置本地推理。结合 encoding/README.md、inference/README.md 与根目录 README.md,即可完成从编码、部署到调优的完整落地。
【免费下载链接】DeepSeek-V4-Pro-0813
相关推荐
DeepSeek-V4-Flash-0731 完全指南:聊天模板编码、DSpark 投机解码与 vLLM/SGLang/本地推理部署实战
DeepSeek V4 Flash 0731 完全指南:聊天模板编码、DSpark 投机解码与 vLLM/SGLang/本地推理部署实战 DeepSeek V4
人工智能大模型基础模型DeepSeekDeepSeek-V4-Pro-0813 本地推理实战指南:权重转换、单机/多机部署与 FP8/FP4 精度切换
DeepSeek V4 Pro 0813 本地推理实战指南:权重转换、单机/多机部署与 FP8/FP4 精度切换 DeepSeek V4 Pro 0813 的官
为什么推理更快?DeepSeek-V4-Flash-Vision-Exp DSpark 投机解码机制深度全解
为什么推理更快?DeepSeek V4 Flash Vision Exp DSpark 投机解码机制深度全解 DeepSeek V4 Flash Vision
大模型基础模型多模态计算机视觉DeepSeek
创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考