--- snapshot-1789684045+++ snapshot-1790665615@@ -2,7 +2,11 @@ Release notes from redis -2026-09-17T15:06:00Z tag:github.com,2008:Repository/156018/8.10.2 2026-09-17T15:08:27Z +2026-09-28T07:34:42Z tag:github.com,2008:Repository/156018/8.12-m02-int 2026-09-28T07:34:42Z + +8.12-m02-int: Explain a test [TIMEOUT] instead of killing it silently (#15879) + +
A hung test run tells us almost nothing today. When no client has made
progress for --timeout seconds, test_server_cron prints the clients'
last reported state, SIGKILLs every server via force_kill_all_servers
and exits; --dump-logs only fires for a failed or excepted test, never
for a timeout. So a 20-minute hang costs a whole CI run and produces a
few lines.
Collect the evidence before tearing the run down:
Crash-report the surviving servers. SIGSEGV makes redis log a stack
trace of every one of its threads plus INFO, the client list and the
config (printCrashReport) and then die. The handler runs on whichever
thread takes the signal, so it works on a server whose event loop is
wedged -- exactly the case we cannot diagnose from outside. kill_server
already resorts to SIGSEGV for the same reason when a server won't exit,
but the timeout path never reaches it. Then print each server's crash
report, starting 10 lines above "REDIS BUG REPORT START" for context (or
the log's tail if there is no report, since that is then the only
evidence). Servers are children of the stuck client, which isn't reaping
them, so a dead one is a zombie that kill -0 still reports alive; wait
on is_running (via ps) instead.
Report where in the test each client stopped, not just its last state.
Crash-reporting the servers usually unblocks a client by itself -- its
connection dies, the error unwinds, and the client's existing top-level
handler reports $::errorInfo, a Tcl stack trace naming the exact line --
so collect that first. A client stuck on something else is poked with
SIGUSR1, which it turns into a Tcl error with Tclx's "signal error":
that interrupts a blocking read, a long "after" and a polling loop
alike. Tclx is optional; without it a timeout simply reports no client
stack trace.
Order matters here: the servers must be collected first, because
unblocking a client makes start_server kill the very servers we wanted a
report from.
Also: read_from_test_client threw "expected non-negative integer" once a
reporting client exited, because we now pump the event loop while it
does.
Note
Low Risk
Changes are limited to the Tcl test harness timeout path; no
production server or runtime behavior is affected.
Overview
When the suite hits --timeout (no client progress), it no longer
tears down immediately after printing each client’s last task. The test
server collects diagnostics first, then kills clients and servers as
before.
Server evidence: Still-running Redis instances from
::active_servers get SIGCONT (if stopped) and SIGSEGV so they
write a full crash report even when the event loop is wedged. Logs are
located under tests/tmp by pid, then dump_crash_report prints
the tail around REDIS BUG REPORT START (or the last 256KB if there is
no report). is_running treats zombies as dead so waits don’t hang
on unreaped children.
Client evidence: Clients send their OS pid on ready and
optionally advertise sigusr1-trace when Tclx is available. After
server dumps, the server waits briefly for natural exception/err
unwinds, then sends SIGUSR1 to remaining clients so Tclx turns it
into a stack trace. ::in_timeout_report suppresses re-entrant
timeout cron, avoids fatal handling of those packets, and fixes
read_from_test_client when a client disconnects mid-report
(invalid length no longer spins or crashes the handler).
Order: Server crash collection runs before client stack traces
so unblocking a client doesn’t tear down servers before their reports
are captured.
Reviewed by Cursor Bugbot for commit
5927e7e. Bugbot is set up for automated
code reviews on this repo. Configure
here.
Co-authored-by: Claude Opus 5.5 (1M context) noreply@anthropic.com
oranagra tag:github.com,2008:Repository/156018/8.10.2 2026-09-17T15:08:27Z 8.10.2 @@ -38,8 +42,4 @@ 8.6.6 -Update urgency: SECURITY: There are security fixes in the release.
CMSketch RDB loading may lead to heap OOB writeSORT, GEORADIUS/GEORADIUSBYMEMBER and XREAD/XREADGROUP: the keys validated by ACL could differ from the keys the command actually accessesSLOT_INFO slot id causes memory corruption during RDB loading, which may lead to Remote Code ExecutionVREM mutates the HNSW graph while background VSIM threads are still runninghnsw_search() return was treated as a huge unsigned count, reading past the end of the result arraysUpdate urgency: SECURITY: There are security fixes in the release.
CMSketch RDB loading may lead to heap OOB writeSORT, GEORADIUS/GEORADIUSBYMEMBER and XREAD/XREADGROUP: the keys validated by ACL could differ from the keys the command actually accessesargv access during key extraction when checking ACL permissions of a KEYNUM keyspec command (e.g. EVAL) with wrong aritySLOT_INFO slot id causes memory corruption during RDB loading, which may lead to Remote Code ExecutionVREM mutates the HNSW graph while background VSIM threads are still runninghnsw_search() return was treated as a huge unsigned count, reading past the end of the result arraysUpdate urgency: SECURITY: There are security fixes in the release.
CMSketch RDB loading may lead to heap OOB writeSORT, GEORADIUS/GEORADIUSBYMEMBER and XREAD/XREADGROUP: the keys validated by ACL could differ from the keys the command actually accessesSLOT_INFO slot id causes memory corruption during RDB loading, which may lead to Remote Code ExecutionVREM mutates the HNSW graph while background VSIM threads are still runninghnsw_search() return was treated as a huge unsigned count, reading past the end of the result arrays