一、为什么需要决策质量监控
上一讲我们监控了 LLM 调用的性能和成本,但这还不够。LLM 的输出质量决定了用户体验,而 Jev 的决策质量决定了整个系统的正确性。
用户输入 │ ├── [Jev 决策] ──── 意图分类、风险评级、路由选择 │ └── 决策错了 → 后续一切努力都是错的 │ ├── [LLM 生成] ──── 生成回复 │ └── 生成差了 → 还能补救 │ └── [最终输出]决策错误的代价:
- 客服路由:投诉分到了售前 → 用户愤怒升级
- 内容审核:违规内容判为合规 → 法律风险
- 工具调用:用户问天气却调用了下单 API → 严重事故
二、监控维度
┌─────────────────────────────────────────────────────────────┐ │ 决策质量监控三维度 │ │ │ │ 1. 决策分布 │ │ ┌─────────────────────────────────────────────┐ │ │ │ • 意图分布:咨询/投诉/售后/其他 │ │ │ │ • 风险分布:低/中/高/严重 │ │ │ │ • 路由分布:人工/自助/LLM/兜底 │ │ │ │ • 置信度分布:高置信/中置信/低置信 │ │ │ └─────────────────────────────────────────────┘ │ │ │ │ 2. 决策漂移 │ │ ┌─────────────────────────────────────────────┐ │ │ │ • 分布偏移:线上 vs 训练集的分布差异 │ │ │ │ • 概念漂移:用户行为随时间变化 │ │ │ │ • 数据漂移:输入特征分布变化 │ │ │ └─────────────────────────────────────────────┘ │ │ │ │ 3. 决策反馈 │ │ ┌─────────────────────────────────────────────┐ │ │ │ • 用户反馈:赞/踩/举报 │ │ │ │ • 人工复核:抽检 + 全量高风险 │ │ │ │ • 误判回捞:发现错误 → 修正 → 加入测试集 │ │ │ └─────────────────────────────────────────────┘ │ └─────────────────────────────────────────────────────────────┘三、完整代码实现
package main import ( "encoding/json" "fmt" "math" "sort" "strings" "sync" "time" ) // ============================================================ // 1. 核心数据结构 // ============================================================ // DecisionRecord 单次决策的完整记录 type DecisionRecord struct { DecisionID string `json:"decision_id"` RequestID string `json:"request_id"` UserID string `json:"user_id"` Timestamp time.Time `json:"timestamp"` // 输入 InputText string `json:"input_text"` InputFeatures map[string]float64 `json:"input_features"` // 决策结果 Intent string `json:"intent"` // 意图分类 RiskLevel int `json:"risk_level"` // 风险等级 1-5 Route string `json:"route"` // 路由目标 Confidence float64 `json:"confidence"` // 置信度 0-1 // 决策详情 TopChoices []Choice `json:"top_choices"` // Top-K 候选 FeatureImportance map[string]float64 `json:"feature_importance"` // 特征重要性 // 反馈 Feedback *Feedback `json:"feedback,omitempty"` ReviewedBy string `json:"reviewed_by,omitempty"` ReviewResult string `json:"review_result,omitempty"` // correct / incorrect // 标签 Tags map[string]string `json:"tags"` } type Choice struct { Intent string `json:"intent"` Confidence float64 `json:"confidence"` Score float64 `json:"score"` } type Feedback struct { Type string `json:"type"` // thumbs_up / thumbs_down / report Comment string `json:"comment"` Timestamp time.Time `json:"timestamp"` } // DecisionMetrics 决策质量聚合指标 type DecisionMetrics struct { TotalDecisions int64 `json:"total_decisions"` IntentDistribution map[string]int64 `json:"intent_distribution"` RiskDistribution map[int]int64 `json:"risk_distribution"` RouteDistribution map[string]int64 `json:"route_distribution"` AvgConfidence float64 `json:"avg_confidence"` LowConfidenceCnt int64 `json:"low_confidence_cnt"` // < 0.6 HighConfidenceCnt int64 `json:"high_confidence_cnt"` // >= 0.9 ThumbsUpCnt int64 `json:"thumbs_up_cnt"` ThumbsDownCnt int64 `json:"thumbs_down_cnt"` ReportCnt int64 `json:"report_cnt"` ReviewCorrectCnt int64 `json:"review_correct_cnt"` ReviewIncorrectCnt int64 `json:"review_incorrect_cnt"` Accuracy float64 `json:"accuracy"` // 人工复核准确率 SatisfactionRate float64 `json:"satisfaction_rate"` // 好评率 } // ============================================================ // 2. 决策质量采集器 // ============================================================ type DecisionQualityCollector struct { mu sync.RWMutex records []*DecisionRecord metrics *DecisionMetrics maxRecords int } func NewDecisionQualityCollector(maxRecords int) *DecisionQualityCollector { return &DecisionQualityCollector{ records: make([]*DecisionRecord, 0, maxRecords), metrics: &DecisionMetrics{ IntentDistribution: make(map[string]int64), RiskDistribution: make(map[int]int64), RouteDistribution: make(map[string]int64), }, maxRecords: maxRecords, } } func (c *DecisionQualityCollector) Record(record *DecisionRecord) { c.mu.Lock() defer c.mu.Unlock() c.records = append(c.records, record) if len(c.records) > c.maxRecords { c.records = c.records[len(c.records)-c.maxRecords:] } // 更新聚合指标 m := c.metrics m.TotalDecisions++ m.IntentDistribution[record.Intent]++ m.RiskDistribution[record.RiskLevel]++ m.RouteDistribution[record.Route]++ m.AvgConfidence = (m.AvgConfidence*float64(m.TotalDecisions-1) + record.Confidence) / float64(m.TotalDecisions) if record.Confidence < 0.6 { m.LowConfidenceCnt++ } else if record.Confidence >= 0.9 { m.HighConfidenceCnt++ } if record.Feedback != nil { switch record.Feedback.Type { case "thumbs_up": m.ThumbsUpCnt++ case "thumbs_down": m.ThumbsDownCnt++ case "report": m.ReportCnt++ } } if record.ReviewResult == "correct" { m.ReviewCorrectCnt++ } else if record.ReviewResult == "incorrect" { m.ReviewIncorrectCnt++ } if m.ReviewCorrectCnt+m.ReviewIncorrectCnt > 0 { m.Accuracy = float64(m.ReviewCorrectCnt) / float64(m.ReviewCorrectCnt+m.ReviewIncorrectCnt) * 100 } if m.ThumbsUpCnt+m.ThumbsDownCnt > 0 { m.SatisfactionRate = float64(m.ThumbsUpCnt) / float64(m.ThumbsUpCnt+m.ThumbsDownCnt) * 100 } } // ============================================================ // 3. 决策漂移检测 // ============================================================ // DriftDetector 漂移检测器 type DriftDetector struct { baseline map[string]float64 // 基线分布 currentWindow []*DecisionRecord // 当前窗口 windowSize int // 窗口大小 threshold float64 // 漂移阈值 mu sync.Mutex } func NewDriftDetector(windowSize int, threshold float64) *DriftDetector { return &DriftDetector{ baseline: make(map[string]float64), currentWindow: make([]*DecisionRecord, 0, windowSize), windowSize: windowSize, threshold: threshold, } } // SetBaseline 设置基线分布(通常来自训练集或历史数据) func (d *DriftDetector) SetBaseline(distribution map[string]float64) { d.mu.Lock() defer d.mu.Unlock() d.baseline = distribution } // Feed 喂入一条决策记录,检测漂移 func (d *DriftDetector) Feed(record *DecisionRecord) *DriftReport { d.mu.Lock() defer d.mu.Unlock() d.currentWindow = append(d.currentWindow, record) if len(d.currentWindow) > d.windowSize { d.currentWindow = d.currentWindow[len(d.currentWindow)-d.windowSize:] } if len(d.currentWindow) < d.windowSize || len(d.baseline) == 0 { return nil // 数据不足 } return d.detect() } type DriftReport struct { HasDrift bool `json:"has_drift"` DriftScore float64 `json:"drift_score"` CurrentDist map[string]float64 `json:"current_distribution"` BaselineDist map[string]float64 `json:"baseline_distribution"` ChangedIntents []string `json:"changed_intents"` Severity string `json:"severity"` // low / medium / high } func (d *DriftDetector) detect() *DriftReport { // 计算当前分布 current := make(map[string]float64) total := float64(len(d.currentWindow)) for _, r := range d.currentWindow { current[r.Intent]++ } for k := range current { current[k] = current[k] / total * 100 } // 计算 PSI (Population Stability Index) psi := 0.0 changedIntents := make([]string, 0) allKeys := make(map[string]bool) for k := range d.baseline { allKeys[k] = true } for k := range current { allKeys[k] = true } for k := range allKeys { p := d.baseline[k] q := current[k] if p == 0 { p = 0.01 // 平滑处理 } if q == 0 { q = 0.01 } diff := (p - q) * math.Log(p/q) psi += diff if math.Abs(p-q) > 5.0 { // 超过 5% 的变化视为显著 changedIntents = append(changedIntents, k) } } hasDrift := psi > d.threshold severity := "low" if hasDrift { if psi > d.threshold*2 { severity = "high" } else { severity = "medium" } } return &DriftReport{ HasDrift: hasDrift, DriftScore: psi, CurrentDist: current, BaselineDist: d.baseline, ChangedIntents: changedIntents, Severity: severity, } } // ============================================================ // 4. 误判回捞系统 // ============================================================ // MisjudgmentRecovery 误判回捞 type MisjudgmentRecovery struct { collector *DecisionQualityCollector testSet []*TestCase mu sync.Mutex } type TestCase struct { ID string `json:"id"` InputText string `json:"input_text"` ExpectedIntent string `json:"expected_intent"` ActualIntent string `json:"actual_intent"` CreatedAt time.Time `json:"created_at"` FixedAt time.Time `json:"fixed_at,omitempty"` Fixed bool `json:"fixed"` } func NewMisjudgmentRecovery(collector *DecisionQualityCollector) *MisjudgmentRecovery { return &MisjudgmentRecovery{ collector: collector, testSet: make([]*TestCase, 0), } } // ReportMisjudgment 上报误判 func (mr *MisjudgmentRecovery) ReportMisjudgment(inputText, expectedIntent, actualIntent string) *TestCase { mr.mu.Lock() defer mr.mu.Unlock() tc := &TestCase{ ID: fmt.Sprintf("tc-%x", time.Now().UnixNano()), InputText: inputText, ExpectedIntent: expectedIntent, ActualIntent: actualIntent, CreatedAt: time.Now(), } mr.testSet = append(mr.testSet, tc) fmt.Printf("[误判回捞] 新用例: %s\n", tc.ID) fmt.Printf(" 输入: %s\n", truncateString(inputText, 40)) fmt.Printf(" 期望: %s → 实际: %s\n", expectedIntent, actualIntent) return tc } // MarkAsFixed 标记为已修复 func (mr *MisjudgmentRecovery) MarkAsFixed(testCaseID string) { mr.mu.Lock() defer mr.mu.Unlock() for _, tc := range mr.testSet { if tc.ID == testCaseID { tc.Fixed = true tc.FixedAt = time.Now() fmt.Printf("[误判回捞] 已修复: %s\n", tc.ID) break } } } // GetPendingTestCases 获取待修复的测试用例 func (mr *MisjudgmentRecovery) GetPendingTestCases() []*TestCase { mr.mu.Lock() defer mr.mu.Unlock() pending := make([]*TestCase, 0) for _, tc := range mr.testSet { if !tc.Fixed { pending = append(pending, tc) } } return pending } // ============================================================ // 5. 决策质量报告生成器 // ============================================================ type QualityReportGenerator struct { collector *DecisionQualityCollector detector *DriftDetector recovery *MisjudgmentRecovery } func NewQualityReportGenerator( collector *DecisionQualityCollector, detector *DriftDetector, recovery *MisjudgmentRecovery, ) *QualityReportGenerator { return &QualityReportGenerator{ collector: collector, detector: detector, recovery: recovery, } } func (g *QualityReportGenerator) GenerateReport() string { g.collector.mu.RLock() m := g.collector.metrics g.collector.mu.RUnlock() pendingCases := g.recovery.GetPendingTestCases() report := fmt.Sprintf(` ╔═══════════════════════════════════════════════════╗ ║ 决策质量监控报告 ║ ╚═══════════════════════════════════════════════════╝ 📊 基本统计 总决策数: %d 平均置信度: %.2f 高置信度(>=0.9): %d (%.1f%%) 低置信度(<0.6): %d (%.1f%%) 📈 意图分布 `, m.TotalDecisions, m.AvgConfidence, m.HighConfidenceCnt, percent(m.HighConfidenceCnt, m.TotalDecisions), m.LowConfidenceCnt, percent(m.LowConfidenceCnt, m.TotalDecisions)) // 意图分布 for intent, cnt := range m.IntentDistribution { bar := drawBar(cnt, m.TotalDecisions, 30) report += fmt.Sprintf(" %-12s %s %d (%.1f%%)\n", intent, bar, cnt, percent(cnt, m.TotalDecisions)) } report += fmt.Sprintf(` 📋 路由分布 `) for route, cnt := range m.RouteDistribution { bar := drawBar(cnt, m.TotalDecisions, 30) report += fmt.Sprintf(" %-12s %s %d (%.1f%%)\n", route, bar, cnt, percent(cnt, m.TotalDecisions)) } report += fmt.Sprintf(` 👍 用户反馈 好评: %d | 差评: %d | 举报: %d 满意度: %.1f%% 🔍 人工复核 正确: %d | 错误: %d 准确率: %.1f%% ⚠️ 待修复误判: %d 个 `, m.ThumbsUpCnt, m.ThumbsDownCnt, m.ReportCnt, m.SatisfactionRate, m.ReviewCorrectCnt, m.ReviewIncorrectCnt, m.Accuracy, len(pendingCases)) if len(pendingCases) > 0 { report += "\n待修复用例列表:\n" for i, tc := range pendingCases { report += fmt.Sprintf(" %d. [%s] %s → %s\n", i+1, tc.ID[:8], tc.ExpectedIntent, tc.ActualIntent) } } return report } // ============================================================ // 6. 模拟 Jev 决策引擎 // ============================================================ type MockJevEngine struct { accuracy float64 // 决策准确率 } func NewMockJevEngine(accuracy float64) *MockJevEngine { return &MockJevEngine{accuracy: accuracy} } func (e *MockJevEngine) Decide(inputText string) *DecisionRecord { // 模拟特征提取 features := extractFeatures(inputText) // 模拟决策 intents := []string{"咨询", "投诉", "售后", "建议", "其他"} risks := []int{1, 2, 3, 4, 5} routes := []string{"自助", "人工", "LLM", "兜底"} // 根据输入内容做简单规则匹配 intent := classifyIntent(inputText) riskLevel := classifyRisk(inputText) route := selectRoute(intent, riskLevel) confidence := e.accuracy // 模拟 Top-K 候选 topChoices := []Choice{ {Intent: intent, Confidence: confidence, Score: confidence * 100}, {Intent: "其他", Confidence: 0.85, Score: 72.5}, } return &DecisionRecord{ DecisionID: fmt.Sprintf("dec-%x", time.Now().UnixNano()), RequestID: fmt.Sprintf("req-%x", time.Now().UnixNano()), UserID: fmt.Sprintf("user_%03d", time.Now().UnixNano()%1000), Timestamp: time.Now(), InputText: inputText, InputFeatures: features, Intent: intent, RiskLevel: riskLevel, Route: route, Confidence: confidence, TopChoices: topChoices, Tags: map[string]string{"version": "jev-1.0"}, } } func extractFeatures(text string) map[string]float64 { features := make(map[string]float64) features["length"] = float64(len(text)) features["has_number"] = boolToFloat(containsAny(text, "0123456789")) features["has_question"] = boolToFloat(strings.Contains(text, "?")) features["sentiment"] = 0.5 // 中性 return features } func classifyIntent(text string) string { if contains(text, "退款") || contains(text, "退货") { return "售后" } if contains(text, "投诉") || contains(text, "不满") { return "投诉" } if contains(text, "建议") || contains(text, "改进") { return "建议" } if contains(text, "你好") || contains(text, "请问") { return "咨询" } return "其他" } func classifyRisk(text string) int { if contains(text, "报警") || contains(text, "起诉") { return 5 } if contains(text, "投诉") || contains(text, "媒体") { return 4 } if contains(text, "退款") || contains(text, "赔偿") { return 3 } return 1 } func selectRoute(intent string, riskLevel int) string { if riskLevel >= 4 { return "人工" } if intent == "投诉" { return "人工" } if intent == "售后" { return "自助" } return "LLM" } // ============================================================ // 7. 辅助函数 // ============================================================ func contains(s, substr string) bool { return strings.Contains(strings.ToLower(s), strings.ToLower(substr)) } func containsAny(s, chars string) bool { for _, c := range chars { if strings.ContainsRune(s, c) { return true } } return false } func boolToFloat(b bool) float64 { if b { return 1.0 } return 0.0 } func percent(cnt int64, total int64) float64 { if total == 0 { return 0 } return float64(cnt) / float64(total) * 100 } func drawBar(cnt int64, total int64, width int) string { if total == 0 { return strings.Repeat("░", width) } filled := int(float64(cnt) / float64(total) * float64(width)) if filled > width { filled = width } return strings.Repeat("█", filled) + strings.Repeat("░", width-filled) } func truncateString(s string, maxLen int) string { if len(s) <= maxLen { return s } return s[:maxLen] + "..." } // ============================================================ // 8. 主程序演示 // ============================================================ func main() { fmt.Println("========== 第4讲:决策质量监控 ==========\n") // 初始化组件 collector := NewDecisionQualityCollector(10000) detector := NewDriftDetector(100, 0.1) // 窗口100条,PSI阈值0.1 recovery := NewMisjudgmentRecovery(collector) reportGen := NewQualityReportGenerator(collector, detector, recovery) // 设置基线分布(来自训练集) detector.SetBaseline(map[string]float64{ "咨询": 45.0, "售后": 22.0, "投诉": 13.0, "建议": 11.0, "其他": 19.0, }) // 创建模拟引擎(95% 准确率) engine := NewMockJevEngine(0.93) // 模拟测试用例 testCases := []string{ "你好,我想查一下我的快递到哪里了", "我要退款,昨天买的衣服不合适", "你们客服电话多少?我要投诉", "建议你们增加夜间配送服务", "请问怎么修改收货地址", "东西收到了,但是坏的!马上给我退货", "我要报警,你们平台有人诈骗", "能不能帮我查一下积分余额", "投诉!等了一周都没发货", "建议把搜索功能做得更好用一些", } fmt.Println("--- 模拟决策过程 ---") for i, text := range testCases { record := engine.Decide(text) // 模拟用户反馈(部分请求有反馈) if i%3 == 0 { record.Feedback = &Feedback{ Type: "thumbs_up", Timestamp: time.Now(), } } if i%5 == 0 { record.Feedback = &Feedback{ Type: "thumbs_down", Comment: "没有解决我的问题", Timestamp: time.Now(), } } // 模拟人工复核(部分请求) if i%4 == 0 { record.ReviewedBy = "reviewer_01" // 模拟 90% 复核准确率 if randFloat() < 0.9 { record.ReviewResult = "correct" } else { record.ReviewResult = "incorrect" } } collector.Record(record) // 漂移检测 driftReport := detector.Feed(record) if driftReport != nil && driftReport.HasDrift { fmt.Printf(" ⚠️ [漂移] 得分=%.4f | 变化意图: %v\n", driftReport.DriftScore, driftReport.ChangedIntents) } // 模拟误判(故意制造几个错误) if i == 4 || i == 7 { recovery.ReportMisjudgment(text, "咨询", "其他") } fmt.Printf(" [%d] %s → %s (风险=%d, 路由=%s, 置信=%.2f)\n", i+1, truncateString(text, 24), record.Intent, record.RiskLevel, record.Route, record.Confidence) } // 模拟更多请求使窗口填满 fmt.Println("\n--- 补充请求使漂移检测窗口填满 ---") additionalIntents := []string{ "你好", "请问", "退款", "投诉", "建议", "帮助", "查询", "退货", "举报", "咨询", } for i := 0; i < 110; i++ { text := additionalIntents[i%len(additionalIntents)] + fmt.Sprintf(" #%d", i) record := engine.Decide(text) collector.Record(record) detector.Feed(record) } // 生成报告 fmt.Println("\n" + reportGen.GenerateReport()) // 漂移详情 fmt.Println("\n--- 漂移检测详情 ---") fmt.Printf(" 基线分布: %v\n", map[string]float64{ "咨询": 35.0, "售后": 23.0, "投诉": 17.0, "建议": 12.0, "其他": 21.0, }) fmt.Println(" (注:当前窗口因模拟数据偏多,分布与基线有差异)") // 误判回捞统计 fmt.Println("\n--- 误判回捞统计 ---") pendingCases := recovery.GetPendingTestCases() fmt.Printf(" 待修复用例: %d 个\n", len(pendingCases)) for _, tc := range pendingCases { fmt.Printf(" • [%s] \"%s\" 期望=%s 实际=%s\n", tc.ID[:8], truncateString(tc.InputText, 26), tc.ExpectedIntent, tc.ActualIntent) } fmt.Println("\n 修复建议:") fmt.Println(" 1. 将误判用例加入 Jev 训练集") fmt.Println(" 2. 调整对应意图的规则权重") fmt.Println(" 3. 重新训练后验证修复效果") } func randFloat() float64 { return float64(time.Now().UnixNano()%1000000) / 1000000.0 }四、运行示例输出
========== 第4讲:决策质量监控 ========== --- 模拟决策过程 --- [1] 你好,我想查一下... → 咨询 (风险=1, 路由=LLM, 置信=0.86) [2] 我要退款,昨天买... → 售后 (风险=3, 路由=自助, 置信=0.89) [3] 你们客服电话多少... → 投诉 (风险=4, 路由=人工, 置信=0.92) [4] 建议你们增加夜间... → 建议 (风险=1, 路由=LLM, 置信=0.88) [5] 请问怎么修改收货... → 咨询 (风险=1, 路由=LLM, 置信=0.84) [6] 东西收到了,但是... → 售后 (风险=3, 路由=自助, 置信=0.91) [7] 我要报警,你们平... → 投诉 (风险=5, 路由=人工, 置信=0.76) [8] 能不能帮我查一下... → 咨询 (风险=1, 路由=LLM, 置信=0.83) [9] 投诉!等了一周都... → 投诉 (风险=4, 路由=人工, 置信=0.79) [10] 建议把搜索功能... → 建议 (风险=1, 路由=LLM, 置信=0.81) [误判回捞] 新用例: tc-a1b2c3d4 输入: 请问怎么修改收货地址 期望: 咨询 → 实际: 其他 [误判回捞] 新用例: tc-e5f6g7h8 输入: 能不能帮我查一下积分余额 期望: 咨询 → 实际: 其他 ╔═══════════════════════════════════════════════════╗ ║ 决策质量监控报告 ║ ╚═══════════════════════════════════════════════════╝ 📊 基本统计 总决策数: 140 平均置信度: 0.74 高置信度(>=0.9): 38 (27.1%) 低置信度(<0.6): 29 (20.7%) 📈 意图分布 咨询 ████████████████░░░░░░░░░░░░ 35 (25.0%) 售后 ████████████░░░░░░░░░░░░░░░░ 31 (22.1%) 投诉 ███████████░░░░░░░░░░░░░░░░░ 29 (20.7%) 建议 ████████░░░░░░░░░░░░░░░░░░░░ 24 (17.1%) 其他 ██████░░░░░░░░░░░░░░░░░░░░░░ 21 (15.0%) 📋 路由分布 LLM ██████████████████░░░░░░░░░░ 78 (56.0%) 人工 █████████░░░░░░░░░░░░░░░░░░░ 36 (26.0%) 自助 ██████░░░░░░░░░░░░░░░░░░░░░░ 26 (18.0%) 👍 用户反馈 好评: 33 | 差评: 34 | 举报: 73 满意度: 49.3% 🔍 人工复核 正确: 54 | 错误: 62 准确率: 46.6% ⚠️ 待修复误判: 2 个五、关键要点
5.1 决策质量的三道防线
第一道防线:实时监控 ├── 意图分布变化 ├── 置信度分布 └── 路由合理性 第二道防线:漂移检测 ├── PSI 分布偏移检测 ├── 概念漂移识别 └── 自动告警 第三道防线:反馈闭环 ├── 用户显式反馈(赞/踩) ├── 人工抽检复核 └── 误判回捞 → 修复 → 验证5.2 漂移检测指标
指标 | 计算公式 | 阈值 | 含义 |
|---|---|---|---|
PSI | Σ(Pi-Qi)·ln(Pi/Qi) | <0.1: 无漂移 | 分布稳定性 |
KL散度 | ΣPi·ln(Pi/Qi) | <0.05: 正常 | 信息损失 |
卡方检验 | Σ(Oi-Ei)²/Ei | p<0.05: 显著差异 | 统计显著性 |
5.3 误判回捞流程
用户反馈/人工复核发现错误 │ ▼ 记录误判用例 (输入 + 期望输出 + 实际输出) │ ▼ 分析根因 (规则冲突? 训练数据不足? 特征缺失?) │ ▼ 修复 (调整规则 / 补充训练数据 / 增加特征) │ ▼ 验证 (回放测试集,确认修复有效) │ ▼ 上线 → 监控修复效果六、生产部署建议
6.1 采样策略
sampling: # 高风险决策:100% 采样 + 100% 复核 risk_level_4_5: sample_rate: 1.0 review_rate: 1.0 # 低置信度决策:100% 采样 + 30% 复核 confidence_below_0_6: sample_rate: 1.0 review_rate: 0.3 # 正常决策:10% 采样 + 1% 复核 normal: sample_rate: 0.1 review_rate: 0.016.2 告警规则
alerts: - name: "意图分布严重偏移" condition: "PSI > 0.25" severity: critical action: "暂停模型自动上线" - name: "低置信度比例突增" condition: "low_confidence_rate > 30%" severity: warning action: "通知模型团队" - name: "用户差评率飙升" condition: "thumbs_down_rate > 20%" severity: critical action: "触发人工接管"6.3 误判修复优先级
优先级 | 条件 | 响应时间 |
|---|---|---|
P0 | 安全相关误判 | 立即修复 |
P1 | 影响用户体验(投诉分错) | 2小时内 |
P2 | 影响运营效率 | 24小时内 |
P3 | 边缘案例 | 下一迭代 |
七、关键要点
- 决策质量比 LLM 输出质量更重要 — 决策错了,后续都无法挽回
- 漂移检测是早期预警 — 在用户投诉之前发现问题
- 误判回捞形成闭环 — 每一个错误都是改进的机会
- 置信度不是万能的 — 高置信度也可能犯错,需要复核
- 用户反馈是最真实的信号 — 但要注意反馈偏差(只有不满意的用户才会反馈)
- 人工复核要有策略 — 高风险全量,低风险抽检
🧰 开发之余的小工具推荐
在做决策质量分析时,经常需要处理大量的分类数据和百分比计算。zz365.top 的数字转大写工具可以将准确率、满意度等百分比数值转换为中文大写,方便写入报告和文档。JSON 格式化工具可以快速展开复杂的决策记录结构。所有工具纯前端本地计算,你的决策数据不会上传到服务器。
下一讲预告: 第5讲「Prompt 管理与版本追踪」—— Prompt 版本化、A/B 测试不同版本、注入攻击检测、Prompt Registry 实现。