1. 生产环境 keys * 到底卡在哪:一次线上阻塞的复盘
RedisTemplate 里用keys *匹配一批 key,本地测试几十毫秒就返回,上了生产却把整个服务拖住——这个场景我猜不少人都遇到过。核心检索词先摆出来:RedisTemplate scan 替代 keys 渐进式遍历,它解决的问题就是「不阻塞 Redis 的前提下,把符合 pattern 的 key 分批捞出来」。适合谁?正在用 Spring Boot + RedisTemplate、需要按前缀清理缓存、做数据迁移、统计某类 key 数量的后端同学。
keys为什么不安全,得从 Redis 的单线程模型说起。Redis 处理命令是单线程串行的,keys pattern执行时会一次性遍历整个 key 空间,把所有匹配结果攒成一个数组返回。这个遍历过程里,Redis 主线程被占满,其他客户端的请求全部排队等待。key 数量到百万级时,keys可能跑几百毫秒甚至几秒,这段时间整个实例对外表现为「卡死」。很多公司的运维直接在redis.conf里用rename-command KEYS ""把命令禁掉,就是这个原因。
scan的设计思路完全不同。它不保证一次返回全部结果,而是维护一个游标(cursor),每次调用返回一小批 key 和一个新游标,客户端拿着新游标继续请求,直到游标回到0表示遍历结束。每次调用只扫描有限数量的槽位,主线程占用时间可控,不会长时间阻塞。代价是:遍历期间新增或删除的 key 可能被漏掉或重复返回,scan只保证「遍历开始时就在的 key 一定被返回至少一次」。
在 RedisTemplate 里,scan的坑比原生命令更多。网上流传的示例大多长这样:connection.scan(ScanOptions...)拿一个Cursor,然后forEachRemaining一把梭。问题在于connection.scan返回的游标在 Spring Data Redis 的封装里没法手动移动,一次调用就把当前批次消费完了,本质上还是「扫一次」,大 key 空间下依然会拉回海量数据。更隐蔽的坑是游标不关闭——Cursor底层持有连接,不close()连接不会归还连接池,跑十几次之后连接池耗尽,后续请求全部卡在JedisConnectionFactory.getConnection,表现为CLOSE_WAIT堆积、线程WAITING。
这篇就按「问题 → 前置准备 → 可复制配置 → 验证 → 排错 → 收尾」的顺序,把 RedisTemplate 下 scan 的正确姿势、连接池配置、以及如何借助 TaoToken 统一 Key/API 通道做联调验证讲清楚。TaoToken 在这里的角色是提供一个统一的模型/API 调用入口,方便你在写代码时快速验证配置片段、排查请求问题,官网入口放在第二节。
2. TaoToken 前置准备:统一 Key 与 API 通道
在动手改 RedisTemplate 之前,先把联调环境准备好。TaoToken 提供统一的 API 通道,官网地址是 https://taotoken.net/?utm_source=taotoken_aicg_blog_end&utm_medium=csdn&utm_campaign=rewrite&utm_content= ,API 根地址是 https://taotoken.net/api 。它的作用是让你在写 scan 配置、调试连接池参数时,有一个稳定的入口去验证请求链路,不用在多个平台之间来回切。
先说清楚它不是什么:TaoToken 不是 Redis 客户端,也不替代你的 RedisTemplate。它是模型/API 的统一通道,你把它理解成「一个统一的调用入口」就行。真正操作 Redis 的还是 Spring Data Redis + Jedis/Lettuce。
准备步骤分三块。
第一块,拿到 API Key。进入控制台 https://taotoken.net/console?utm_source=taotoken_aicg_blog_end&utm_content=console&utm_campaign=rewrite ,在 API Keys 页面创建一个新 key,复制保存。这个 key 后面在验证请求时会用到。API Keys 直达: https://taotoken.net/api-keys?utm_source=taotoken_aicg_blog_end&utm_content=api-keys&utm_campaign=rewrite 。
第二块,确认 Base URL 和 Model ID。Base URL 用https://taotoken.net/api,Model ID 根据你实际要调用的模型填。这三个要素——Base URL、Key、Model ID——是任何接入场景的「三件套」,缺一个都跑不通。如果你用的是 Claude Code 这类编码工具,配置入口在 https://taotoken.net/claude-code-anthropic?utm_source=taotoken_aicg_blog_end&utm_content=claudecode&utm_campaign=rewrite ;如果是 Coding Plan 长期编码场景,看 https://taotoken.net/coding-plan?utm_source=taotoken_aicg_blog_end&utm_content=coding-plan&utm_campaign=rewrite 。
第三块,验证通道连通。用 curl 发一个最小请求,确认 Base URL 和 Key 没问题:
curl -X POST https://taotoken.net/api/v1/chat/completions \ -H "Authorization: Bearer $TAOTOKEN_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "your-model-id", "messages": [{"role": "user", "content": "ping"}] }'返回里能看到choices字段就说明通道正常。这一步的意义在于:后面你调 RedisTemplate 的 scan 逻辑时,如果怀疑是网络或鉴权问题,可以先用这个请求排除掉通道因素,把问题范围缩小到 Redis 侧。
文档入口放在这里方便查参数: https://taotoken.net/doc?utm_source=taotoken_aicg_blog_end&utm_content=doc&utm_campaign=rewrite 。模型对话调试入口: https://taotoken.net/chat?utm_source=taotoken_aicg_blog_end&utm_content=chat&utm_campaign=rewrite 。
注意:TaoToken 的 Key 只用于 API 通道鉴权,不要和 Redis 的密码混用,也不要写进前端代码。Redis 密码走
RedisStandaloneConfiguration.setPassword,两者是独立的凭证体系。
前置准备做完,你手上应该有:一个可用的 TaoToken API Key、确认过的 Base URL、一个能返回choices的验证请求。接下来进入 RedisTemplate 的 scan 配置。
3. 可复制配置:RedisTemplate scan 与连接池参数
这一节给可直接粘贴的配置片段。分两部分:RedisTemplate 的 scan 封装,以及 Jedis 连接池配置。路径和参数名保持和 Spring Data Redis 一致,你按自己项目的包名调整即可。
先看 scan 的正确封装。核心是用RedisConnection拿到原生连接,通过MultiKeyCommands手动移动游标,并且务必在 finally 里关闭 cursor:
import org.springframework.data.redis.connection.RedisConnection; import org.springframework.data.redis.core.Cursor; import org.springframework.data.redis.core.RedisCallback; import org.springframework.data.redis.core.ScanOptions; import org.springframework.data.redis.core.StringRedisTemplate; import org.springframework.stereotype.Component; import redis.clients.jedis.JedisCommands; import redis.clients.jedis.MultiKeyCommands; import redis.clients.jedis.ScanParams; import redis.clients.jedis.ScanResult; import java.nio.charset.StandardCharsets; import java.util.HashSet; import java.util.Set; @Component public class RedisScanHelper { private final StringRedisTemplate redisTemplate; public RedisScanHelper(StringRedisTemplate redisTemplate) { this.redisTemplate = redisTemplate; } /** * 渐进式遍历,手动移动游标,避免一次性拉回全部 key */ public Set<String> scanKeys(String pattern, int count) { return redisTemplate.execute((RedisCallback<Set<String>>) connection -> { Set<String> result = new HashSet<>(); JedisCommands commands = (JedisCommands) connection.getNativeConnection(); MultiKeyCommands multiKeyCommands = (MultiKeyCommands) commands; ScanParams params = new ScanParams(); params.match(pattern); params.count(count); String cursor = "0"; do { ScanResult<String> scanResult = multiKeyCommands.scan(cursor, params); result.addAll(scanResult.getResult()); cursor = scanResult.getStringCursor(); } while (!"0".equals(cursor)); return result; }); } /** * 基于 Cursor 的遍历,注意 try-with-resources 自动关闭 */ public void scanWithCursor(String pattern, java.util.function.Consumer<byte[]> consumer) { redisTemplate.execute((RedisCallback<Void>) connection -> { try (Cursor<byte[]> cursor = connection.scan( ScanOptions.scanOptions().match(pattern).count(1000).build())) { cursor.forEachRemaining(consumer); } catch (Exception e) { throw new RuntimeException("scan failed", e); } return null; }); } }关键点说明:multiKeyCommands.scan(cursor, params)每次返回一批结果和新游标,循环直到游标为"0"。count是提示值不是硬限制,Redis 可能返回多于或少于 count 的数量,别拿它当精确分页。scanWithCursor用 try-with-resources 保证Cursor关闭,这是避免连接泄漏的关键。
再看连接池配置。默认JedisPoolConfig的maxTotal是 8,scan 操作如果游标不关,8 个连接很快被占满。配置片段如下:
import org.springframework.context.annotation.Bean; import org.springframework.context.annotation.Configuration; import org.springframework.data.redis.connection.RedisConnectionFactory; import org.springframework.data.redis.connection.RedisStandaloneConfiguration; import org.springframework.data.redis.connection.jedis.JedisClientConfiguration; import org.springframework.data.redis.connection.jedis.JedisConnectionFactory; import org.springframework.data.redis.connection.RedisPassword; import redis.clients.jedis.JedisPoolConfig; import java.time.Duration; @Configuration public class RedisConfig { @Bean public RedisConnectionFactory redisConnectionFactory() { RedisStandaloneConfiguration standalone = new RedisStandaloneConfiguration(); standalone.setHostName("127.0.0.1"); standalone.setPort(6379); standalone.setPassword(RedisPassword.of("your-redis-password")); JedisPoolConfig poolConfig = new JedisPoolConfig(); poolConfig.setMaxTotal(32); poolConfig.setMaxIdle(16); poolConfig.setMinIdle(4); poolConfig.setMaxWaitMillis(5000); poolConfig.setTestOnBorrow(true); JedisClientConfiguration clientConfig = JedisClientConfiguration.builder() .readTimeout(Duration.ofSeconds(30)) .connectTimeout(Duration.ofSeconds(5)) .usePooling() .poolConfig(poolConfig) .build(); return new JedisConnectionFactory(standalone, clientConfig); } }maxTotal调到 32 是给 scan 留余量,maxWaitMillis设 5000 表示 5 秒拿不到连接就抛Could not get a resource from the pool,比无限等待好排查。readTimeout30 秒是防止某个 scan 卡住时连接被永久占用。
如果你用 Lettuce 而不是 Jedis,getNativeConnection()返回的是io.lettuce.core.api.StatefulRedisConnection,scan 的 API 不一样,需要走connection.scan(ScanArgs)拿KeyScanCursor。上面这套是 Jedis 路径,别混用。
提示:
ScanOptions.scanOptions().count(Long.MAX_VALUE)这种写法等于一次性拉全部,和 keys 没区别,别用。count 给 500 到 2000 之间比较稳。
配置片段就这些。下一节验证请求和成功结果。
4. 验证请求与成功结果:scan 替换前后对比
配置写完,得验证两件事:scan 能正确返回 key,以及它确实不阻塞。先写一个对比脚本。
准备测试数据,往 Redis 里塞 10 万个带前缀的 key:
redis-cli -h 127.0.0.1 -p 6379 -a your-redis-password > for i in $(seq 1 100000); do redis-cli set "user:profile:$i" "v$i"; done生产环境别这么干,用 pipeline 批量写。数据准备好后,写一个 Spring Boot 测试类对比 keys 和 scan:
import org.junit.jupiter.api.Test; import org.springframework.beans.factory.annotation.Autowired; import org.springframework.boot.test.context.SpringBootTest; import org.springframework.data.redis.core.StringRedisTemplate; import java.util.Set; @SpringBootTest public class ScanVsKeysTest { @Autowired private StringRedisTemplate redisTemplate; @Autowired private RedisScanHelper scanHelper; @Test public void compareKeysAndScan() { long t1 = System.currentTimeMillis(); Set<String> keysResult = redisTemplate.keys("user:profile:*"); long t2 = System.currentTimeMillis(); System.out.println("keys 耗时: " + (t2 - t1) + "ms, 数量: " + keysResult.size()); long t3 = System.currentTimeMillis(); Set<String> scanResult = scanHelper.scanKeys("user:profile:*", 1000); long t4 = System.currentTimeMillis(); System.out.println("scan 耗时: " + (t4 - t3) + "ms, 数量: " + scanResult.size()); } }实测下来,10 万 key 的场景,keys单次耗时可能到 200ms 以上,期间其他请求全部排队;scan总耗时可能更长(因为要多次往返),但每次调用只占几毫秒,其他请求能穿插执行。这就是「总耗时换阻塞时间」的取舍。
验证 scan 不阻塞,开两个终端。终端 A 持续发PING:
redis-cli -h 127.0.0.1 -p 6379 -a your-redis-password --latency终端 B 触发 scan 遍历。观察终端 A 的延迟曲线,scan 期间延迟应该保持在低位波动,不会出现尖峰。换成keys再跑一次,终端 A 会看到明显延迟尖峰。
成功结果长这样:scan 返回的 key 数量和 keys 一致(遍历期间无写入的前提下),日志里每次 scan 调用的耗时在个位数毫秒,redis-cli info clients显示连接数稳定,没有持续增长。
再验证一下连接不泄漏。跑 100 次 scan 后执行:
redis-cli -h 127.0.0.1 -p 6379 -a your-redis-password client list | grep -c "cmd=scan"如果 cursor 正确关闭,这个数字应该接近 0(scan 命令执行完连接就归还了)。如果数字持续增长到几十上百,说明 cursor 没关,连接泄漏了。
TaoToken 通道的验证在这里的作用是:当你怀疑是网络层问题时,用第二节的 curl 请求确认通道正常,把问题锁定在 Redis 侧。模型对话入口可以快速发一个请求看返回: https://taotoken.net/chat?utm_source=taotoken_aicg_blog_end&utm_content=chat&utm_campaign=rewrite 。
5. 本篇常见错排查:401、连接池耗尽与游标泄漏
这一节对照真实报错逐个排。scan 相关的坑集中在连接管理和游标处理上。
报错一:Could not get a resource from the pool
这是连接池耗尽的典型报错。触发路径:scan 的Cursor没关闭,每次调用占一个连接不归还,跑满maxTotal后新请求拿不到连接,maxWaitMillis超时抛这个错。排查步骤:先redis-cli client list看cmd=scan的连接数,如果持续增长就是泄漏。修复:确保Cursor在 try-with-resources 里,或者手动cursor.close()放在 finally。同时把maxTotal调大只是缓解,根治还是关游标。
报错二:401 Unauthorized
这个通常出现在 TaoToken 通道验证时,不是 Redis 的问题。原因:API Key 没带、带错、或者 Base URL 写成了https://taotoken.net(少了/api)。检查三件套:Base URL 用https://taotoken.net/api,Key 从控制台复制完整,Model ID 填对。用第二节的 curl 命令复现,看返回体里的错误信息。
报错三:local proxy failed/ 连接超时
这个报错说明请求没到达目标。可能原因:本地网络策略、Base URL 拼错、或者请求发到了错误的端口。先确认curl https://taotoken.net/api能通,再检查代码里的地址。如果是 Redis 侧的超时,检查readTimeout和connectTimeout配置,以及 Redis 实例的网络可达性。
报错四:reading choices相关解析错误
这个出现在解析 API 返回时,通常是返回体不是预期的 JSON 结构。原因可能是请求被中间层拦截返回了 HTML 错误页,或者 Model ID 不存在导致返回了错误结构。用 curl 直接看原始返回,确认是 JSON 且含choices字段。
报错五:OAuth相关鉴权失败
如果用的是 Claude Code 或类似工具接入,出现 OAuth 报错说明鉴权流程没走完。检查工具配置里的 Base URL、Key、Model ID 三件套是否齐全。Claude Code 的配置入口在 https://taotoken.net/claude-code-anthropic?utm_source=taotoken_aicg_blog_end&utm_content=claudecode&utm_campaign=rewrite ,按文档把三件套填全。
报错六:scan 返回结果重复或遗漏
这不是报错,是 scan 的语义特性。遍历期间有写入或删除时,重复和遗漏是正常的。如果你的业务要求强一致,scan 不适用,得换其他方案(比如维护一个索引 key)。如果只是清理缓存、统计数量,scan 够用。
报错七:CLOSE_WAIT堆积
netstat -an | grep 6379看到大量CLOSE_WAIT,说明应用侧没有正确关闭 socket。根因还是连接没归还连接池。检查 cursor 关闭逻辑,以及RedisConnectionUtils.releaseConnection是否被正确调用。用redisTemplate.execute的回调方式,Spring 会自动释放连接;如果你手动getConnection(),必须手动释放。
排错时如果怀疑是通道问题,用 API Keys 页面重新生成一个 key 测试: https://taotoken.net/api-keys?utm_source=taotoken_aicg_blog_end&utm_content=api-keys&utm_campaign=rewrite 。接入文档在 https://taotoken.net/doc?utm_source=taotoken_aicg_blog_end&utm_content=doc&utm_campaign=rewrite 。
6. 收尾:把 scan 用对的关键习惯
写到这里,核心的东西都给了。最后说几个我踩过的坑,你照着做能少走弯路。
第一个习惯:任何Cursor都必须关。不管是connection.scan返回的,还是opsForHash().scan返回的,try-with-resources 包起来。这是 scan 相关故障里最高频的根因。
第二个习惯:count别给Long.MAX_VALUE。给 500 到 2000,让 Redis 分批返回。给最大值等于放弃 scan 的渐进优势。
第三个习惯:连接池maxTotal别用默认的 8。scan 是长连接操作,8 个连接在高并发下不够用。调到 32 起步,配合maxWaitMillis快速失败。
第四个习惯:生产环境用rename-command禁掉keys。代码里也别留redisTemplate.keys()的调用,用 scan 封装替代。
第五个习惯:验证通道和验证 Redis 分开做。TaoToken 的 curl 请求确认通道,redis-cli --latency确认 Redis 不阻塞,两个独立验证,出问题时能快速定位是哪一侧。
长期做编码和 Agent 场景的话,Coding Plan 入口在 https://taotoken.net/coding-plan?utm_source=taotoken_aicg_blog_end&utm_content=coding-plan&utm_campaign=rewrite ,模型对话调试在 https://taotoken.net/chat?utm_source=taotoken_aicg_blog_end&utm_content=chat&utm_campaign=rewrite 。把 scan 封装和连接池配置落到项目里,跑一遍对比脚本,看延迟曲线,你就知道 keys 该退休了。