前言
「网站变慢了」「偶尔 502」「用户说提交没反应」——这类问题最麻烦的地方不是修,而是定位:它们往往无法在本地复现,也没有明确的报错页面。可用的证据几乎全在日志里,但它们散落在四五个地方:nginx 的 access.log 和 error.log、PHP-FPM 的 error log 与 slowlog、PHP 自己的 error_log、以及应用层日志。
先说明版本问题:「日志分析」本身不是某个 PHP 版本的特有功能,PHP 5.x 时代也是这么做的。但 PHP 8.1 确实给日志排查带来了一批新的、容易被误判的信号——它引入了若干弃用(deprecation)警告,比如浮点数到整数的隐式转换、把null传给内部函数的非空参数、以及实现了Serializable之类的内部接口却没有声明返回类型。这些警告在日志里会大量刷屏,很多人的第一反应是「把error_reporting调低」,结果把真正的问题一起关掉了。PHP 8.1 也为排查本身提供了便利:枚举(enum)、readonly属性、never返回类型、#[\ReturnTypeWillChange]属性都是这个版本引入的,写分析脚本时正好用得上。
本文讲三件事:日志怎么串起来看、在 PHP 8.1 环境下哪些日志噪音必须识别、以及一个能直接跑起来的日志聚合脚本。
一、先把证据按层摆好
排查任何「慢/错/丢」的问题,第一步都是确认在哪一层出了问题。下面这张表建议直接贴进项目文档。
| 日志 | 位置(以典型部署为例) | 回答什么问题 | 关键字段 |
|---|---|---|---|
| nginx access.log | 自定log_format指定 | 请求有没有到达、状态码分布、耗时 | $status、$request_time、$request_id |
| nginx error.log | nginx 配置指定 | 反向代理层错误、上游超时 | upstream timed out、connect() failed |
| PHP-FPM error log | error_log指令 | PHP 进程级错误、worker 崩了 | pool www、child exited on signal |
| PHP-FPM slowlog | slowlog指令 +request_slowlog_timeout | 哪个函数的调用栈卡住了 | 慢请求的完整 PHP 调用栈 |
| PHP error_log | php.ini的error_log | 脚本级 Warning / Fatal | 文件、行号、消息 |
| 应用日志 | 项目自己写 | 业务上下文(订单号、用户 ID) | 自己定义,建议统一 JSON |
把 nginx 的$request_id贯穿到 PHP 侧,是性价比最高的一步。nginx 内置了这个变量,不需要额外模块:
log_format main '$remote_addr - $remote_user [$time_local] "$request" ' '$status $body_bytes_sent "$http_referer" "$http_user_agent" ' '$request_time $upstream_response_time $request_id'; server { access_log /var/log/nginx/access.log main; location ~ \.php$ { fastcgi_pass 127.0.0.1:9000; # 把 request_id 透传给 PHP,这样日志可以按同一个 ID 串起来 fastcgi_param HTTP_X_REQUEST_ID $request_id; include fastcgi_params; } }PHP 侧在引导文件里读出来存成常量,之后每条日志都带上它:
define('REQUEST_ID', $_SERVER['HTTP_X_REQUEST_ID'] ?? bin2hex(random_bytes(8)));这样出问题时,只要用户报一个 ID,你就能同时捞出 nginx 那一行、PHP-FPM 的慢查询栈、以及应用日志里的业务上下文。
二、PHP 8.1 环境下必须认识的几类日志噪音
PHP 8.1 的弃用警告不是「错误」,但它们是升级到 8.2、8.3 时变成真错误的种子。在日志里看到下面这些,不要第一反应就去降error_reporting:
| 日志片段 | 含义 | 处理方向 |
|---|---|---|
Deprecated: Implicit conversion from float ... to int loses precision | 浮点数被隐式截断成整数 | 显式写(int)、intdiv()或round(),让意图可见 |
Deprecated: strlen(): Passing null to parameter #1 ($string) of type string | 把null传给了内部函数的非空参数 | 调用前用?? ''兜底,或修正上游的数据来源 |
Deprecated: Return type of X::y() should either be compatible ... or the #[\ReturnTypeWillChange] attribute | 实现内部接口(如ArrayAccess、Iterator)时缺返回类型 | 补上正确的返回类型;一时改不完可以先用#[\ReturnTypeWillChange]属性压住 |
Deprecated: Automatic conversion of false to array is deprecated | 把false当数组用,自动变成了数组 | 初始化时就明确写成[],别依赖自动转换 |
这几类警告的共同特征是数量大但不致命,所以看起来像噪音。真实的排查经验是:先把它们的数量统计出来,按「文件 + 行号」聚合去重——通常几百条日志只对应三五个真实的代码位置。
三、动手写一个聚合脚本
下面这个脚本自身包含一份示例日志,直接运行就能看到分析结果,不需要任何外部依赖。用到enum与readonly属性,都是 PHP 8.1 引入的。
<?php declare(strict_types=1); /** * analyze_log.php —— 聚合 nginx access.log,找出真正的问题请求 * 最低 PHP 8.1(使用 enum 与 readonly 属性) * 用法:php analyze_log.php [日志文件] */ enum Level: string { case Info = 'INFO'; case Warn = 'WARN'; case Error = 'ERROR'; } final class LogLine { public function __construct( public readonly string $ip, public readonly string $uri, public readonly int $status, public readonly float $requestTime, public readonly string $requestId, ) {} } /** 生成一份示例日志,字段顺序与第一节的 log_format 一致 */ function ensure_sample_log(string $path): string { if (is_file($path)) { return $path; } $lines = [ '10.0.0.5 - - [29/Sep/2026 10:00:01 +0800] "GET /api/order?page=1 HTTP/1.1" 200 812 "-" "curl/8.0" 0.031 0.028 5f1c9a2b', '10.0.0.7 - - [29/Sep/2026 10:00:02 +0800] "GET /api/order?page=2 HTTP/1.1" 500 120 "-" "curl/8.0" 3.402 3.398 8ab77e10', '10.0.0.9 - - [29/Sep/2026 10:00:03 +0800] "POST /api/pay HTTP/1.1" 502 157 "-" "curl/8.0" 30.001 - c41d0f88', '10.0.0.5 - - [29/Sep/2026 10:00:04 +0800] "GET /api/order?page=3 HTTP/1.1" 500 120 "-" "curl/8.0" 3.188 3.180 9de0aa31', '10.0.0.11 - - [29/Sep/2026 10:00:05 +0800] "GET /api/order?page=1 HTTP/1.1" 200 812 "-" "curl/8.0" 0.045 0.041 77b3c5ee', ]; file_put_contents($path, implode(PHP_EOL, $lines) . PHP_EOL); return $path; } /** * 解析一行日志。注意 " 与 [] 的分段:$request 里可能包含空格, * 所以必须按引号整体匹配,不能用 \S+ 逐段切。 */ function parse_line(string $line): ?LogLine { // 分组顺序:1=ip 2=时间 3=方法 4=URI 5=状态码 6=字节数 7=UA 8=request_time 9=upstream 10=request_id $pattern = '/^(\S+) \S+ \S+ \[([^\]]+)\] ' . '"([A-Z]+) (\S+) [^"]*" ' . '(\d{3}) (\d+|-) "\S*" "([^"]*)" ' . '([0-9.]+|-) ([0-9.]+|-) ([0-9a-f]+)$/'; if (preg_match($pattern, $line, $m) !== 1) { return null; // 格式不认识的行走这条分支,不要静默丢弃,要计数 } return new LogLine( ip: $m[1], uri: $m[4], status: (int)$m[5], requestTime: (float)$m[8], requestId: $m[10], ); } /** 判断日志级别,用于给聚合结果分层展示 */ function level_of(int $status): Level { return match (true) { $status >= 500 => Level::Error, $status >= 400 => Level::Warn, default => Level::Info, }; } // ---------- 主流程 ---------- $logFile = $argv[1] ?? 'access.log'; $logFile = ensure_sample_log($logFile); $byStatus = []; // 状态码 => 次数 $byUri = []; // 5xx 的 URI => 次数 $slow = []; // 慢请求列表 $badLines = 0; $total = 0; $handle = fopen($logFile, 'rb'); // 流式读取,几百 MB 也不会吃内存 while (($line = fgets($handle)) !== false) { $total++; $parsed = parse_line(rtrim($line, "\r\n")); if ($parsed === null) { $badLines++; continue; } $byStatus[$parsed->status] = ($byStatus[$parsed->status] ?? 0) + 1; if ($parsed->status >= 500) { $byUri[$parsed->uri] = ($byUri[$parsed->uri] ?? 0) + 1; } $slow[] = $parsed; } fclose($handle); printf("总行数 %d,无法解析 %d 行%s", $total, $badLines, PHP_EOL); ksort($byStatus); echo '状态码分布:', PHP_EOL; foreach ($byStatus as $status => $count) { printf(" [%s] %d %d 次%s", level_of($status)->value, $status, $count, PHP_EOL); } arsort($byUri); echo '5xx 命中的 URI:', PHP_EOL; foreach ($byUri as $uri => $count) { printf(" %d 次 %s%s", $count, $uri, PHP_EOL); } // 按耗时排序取前 3 名,并给出 p95 usort($slow, static fn(LogLine $a, LogLine $b): int => $b->requestTime <=> $a->requestTime); echo '最慢的 3 个请求:', PHP_EOL; foreach (array_slice($slow, 0, 3) as $item) { printf(" %.3fs %d %s request_id=%s%s", $item->requestTime, $item->status, $item->uri, $item->requestId, PHP_EOL); } $durations = array_map(static fn(LogLine $l): float => $l->requestTime, $slow); sort($durations); $p95 = $durations === [] ? 0.0 : $durations[(int)floor(0.95 * (count($durations) - 1))]; printf("p95 请求耗时:%.3fs%s", $p95, PHP_EOL);总行数 5,无法解析 0 行 状态码分布: [INFO] 200 2 次 [ERROR] 500 2 次 [ERROR] 502 1 次 5xx 命中的 URI: 2 次 /api/order?page=2 2 次 /api/order?page=3 1 次 /api/pay 最慢的 3 个请求: 30.001s 502 /api/pay request_id=c41d0f88 3.402s 500 /api/order?page=2 request_id=8ab77e10 3.188s 500 /api/order?page=3 request_id=9de0aa31 p95 请求耗时:30.001s拿到request_id=c41d0f88之后,直接去 grep 三个日志文件,就能看到 nginx 报的上游超时、PHP-FPM 里的慢调用栈、以及应用日志里那次支付请求走到了哪一步。这个脚本是排查的起点,不是终点。
常见坑点
- ❌ 用
file()或file_get_contents()一次性读几百 MB 的 access.log。
✅ 用fopen()+fgets()流式处理,内存占用与文件大小无关;需要跨行状态时再用SplFileObject。
- ❌ 正则用
.*贪婪匹配,或者对$request字段用\S+逐段切。
✅ 请求行和 User-Agent 里都可能包含空格,必须按"成对匹配,并在末尾加$锚定,否则半截日志会被误判为有效行。
- ❌ 拿日志里的
$request_time直接和 PHP 里microtime()测出的耗时比较。
✅$request_time是从 nginx 收到首字节起的全链路耗时,包含了排队与网络;两者口径不同,比较前先确认测的是同一段。
- ❌ 看到 4xx 一律当「用户自己输错了」,不纳入统计。
✅ 4xx 突增往往是接口参数变更、鉴权失效或爬虫,要按 URI 聚合看趋势。
- ❌ 排查慢请求只翻 access.log,不看 PHP-FPM 的 slowlog。
✅ 配置request_slowlog_timeout和slowlog,慢请求会自动落一份 PHP 调用栈,能直接指到具体函数。
- ❌ 把用户提交的内容原样写进日志,包含换行符。
✅ 写日志前替换掉\r\n,否则攻击者可以用一个伪造的「日志行」污染分析结果,也就是日志注入(log injection)。
- ❌ 因为 PHP 8.1 的 Deprecated 警告太多,直接把
error_reporting调成只报 Fatal。
✅ 保持E_ALL,把 Deprecated 按「文件:行号」聚合去重后逐条修;升级到更高版本前清干净是必经步骤。
- ❌ 用
shell_exec("grep $keyword access.log")处理用户传入的筛选条件。
✅ 用escapeshellarg()转义,或干脆在 PHP 里读文件再用str_contains()过滤;拼接命令字符串就是命令注入。
总结
| 现象 | 优先看哪里 | 典型线索 |
|---|---|---|
| 全部请求都慢 | nginx access.log 的$request_time分布 | 是否所有 URI 都慢,判断是入口层还是业务层 |
| 个别接口慢 | PHP-FPM slowlog | 调用栈停在哪个函数,是 DB 还是外部 HTTP |
| 间歇 502 | nginx error.log + PHP-FPM 日志 | upstream timed out或child exited,看是超时还是进程崩 |
| 接口偶发 500 | PHP error_log + 应用日志 | 按request_id串联,比按时间戳猜可靠得多 |
| 日志里出现大量 Deprecated | PHP error_log | 按文件行号聚合,属于升级前的必修项 |
分析日志的关键不是工具多强,而是先把各层的证据用同一个请求 ID 串起来,再让统计代替猜测。PHP 8.1 带来的那批 Deprecated 也别急着关掉——它们数量大但位置集中,按文件行号聚合去重之后,通常一次就能修完一大半。