diff --git a/docs/README.md b/docs/README.md index 463a033..c5f13df 100644 --- a/docs/README.md +++ b/docs/README.md @@ -24,12 +24,9 @@ ```bash # 单样本求解(base64 或文件路径) -.venv/bin/python src/solve.py <大图> +.venv/bin/python src/captcha_solver.py <大图> -# 批量求解 + 可视化(输出 src/out/L/) -.venv/bin/python src/method_l_shape.py - -# 结果验证(叠加轮廓图,输出 src/out/L/verify/) +# 结果验证(叠加轮廓图) .venv/bin/python src/verify_result.py ``` diff --git a/docs/solution.md b/docs/solution.md index 414448b..0a012be 100644 --- a/docs/solution.md +++ b/docs/solution.md @@ -22,7 +22,7 @@ 3. **照片纹理会产生大量虚假贴合**:纯倒角距离会被草地等密纹理区域击败。 必须叠加方向约束与洞内部特征。 -## 最终算法(`src/method_l_shape.py`) +## 最终算法(`src/captcha_solver.py` 内部实现) ```text 粗扫:尺度 0.85–1.15(步长 0.025)× 全图位置 @@ -65,14 +65,14 @@ Canny 阈值固定 `lo = max(40, 0.5×中位灰度)`、`hi = 3×lo`。 | d7a7f80c | SIFT + RANSAC | (167.8, 146.4) | (171, 146) | ~4px | | 444d0ea7 | 闭合轮廓分析(Hu≈0.003) | 左洞 rot≈0° | (183, 160) rot=0° | 一致 | -## API(`src/solve.py`) +## API(`src/captcha_solver.py`) ```python -from solve import solve +from captcha_solver import solve_slide result = solve(big_image_b64, mark_image_b64) # 支持 dataURI 前缀 ``` -命令行:`python src/solve.py <大图路径或base64> ` +命令行:`python src/captcha_solver.py <大图路径或base64> ` 返回: @@ -201,7 +201,6 @@ s±0.0375(步长 0.0125)重扫 + ±6px 窗口局部最优,取 rotation_sca ## 复现 ```bash -.venv/bin/python src/method_l_shape.py # 批量求解 + 可视化 .venv/bin/python src/verify_result.py # 叠加轮廓验证图 -.venv/bin/python src/solve.py <大图> # 单样本 API +.venv/bin/python src/captcha_solver.py <大图> # 单样本 API ``` diff --git a/pyrightconfig.json b/pyrightconfig.json new file mode 100644 index 0000000..55fc51b --- /dev/null +++ b/pyrightconfig.json @@ -0,0 +1,6 @@ +{ + "include": ["src", "try"], + "extraPaths": ["src"], + "reportMissingImports": "warning", + "reportMissingModuleSource": "none" +} diff --git a/src/README_3rd.md b/src/README_3rd.md new file mode 100644 index 0000000..94d863c --- /dev/null +++ b/src/README_3rd.md @@ -0,0 +1,264 @@ +# 拖动重叠验证码求解 —— 第三方对接文档 + +## 1. 功能概述 + +输入验证码的**背景大图**与 **mark 小图**(各一张),返回把 mark 拖动重叠到正确目标位置所需的**水平拖动像素**。纯本地图像计算,无任何网络请求。 + +典型场景:抖音系「拖动滑块完成拼图」验证码(`/captcha/get` 返回 `question.url1`(大图 552×344)与 `question.url2`(mark 110×110 PNG 含透明通道))。 + +## 2. 环境要求 + +- Python ≥ 3.10 +- 依赖:`opencv-python`、`numpy`(无其它依赖,全部离线计算) +- 单次求解耗时约 4~7 秒(CPU),无 GPU 需求 + +```bash +pip install opencv-python numpy +``` + +## 3. 接入步骤 + +只依赖 `src/` 下两个文件: + +``` +src/ +└── captcha_solver.py # 对外 API(自包含单文件,唯一交付物) +``` + +### 3.1 最小示例 + +```python +import base64 +import requests +from captcha_solver import solve_slide + +# 1) 从 /captcha/get 响应中取图与 tip_y +q = captcha_get_response["data"]["question"] +big_b64 = base64.b64encode(requests.get(q["url1"]).content).decode() +mark_b64 = base64.b64encode(requests.get(q["url2"]).content).decode() + +# 2) 求解(tip_y 强烈建议传入) +result = solve_slide(big_b64, mark_b64, tip_y=q["tip_y"]) + +if not result["ok"]: + raise RuntimeError(result["message"]) # 解码失败/未找到目标 + +# 3) 拿到拖动像素(原图坐标系) +print(result["distance_px"], result["x"], result["y"], result["quality"], result["confidence"]) +``` + +### 3.2 命令行联调(不写代码先验证效果) + +```bash +python captcha_solver.py big.jpeg mark.png --tip-y 48 +``` + +返回 JSON,字段含义同下。 + +## 4. 返回字段 + +| 字段 | 类型 | 说明 | +| --- | --- | --- | +| `ok` | bool | 求解是否成功 | +| `distance_px` | int | **需要拖动的水平像素(原图 552 坐标系)**,见 §5 换算 | +| `x`, `y` | int | mark 画布左上角在原图中的目标位置(`distance_px == x`) | +| `scale` | float | 目标区域相对 mark 的缩放(0.7~1.1,参考信息) | +| `rot` | float | 目标区域旋转角(±90°,参考信息) | +| `quality` | float | 置信分数,**越小越可信**;经验阈值见 §6 | +| `confidence` | str | `"OK"` 或 `"AMBIGUOUS"`;AMBIGUOUS 建议弃题重刷 | +| `alternatives` | list | 其余候选位置(最多 5 个),可用于换题决策或交叉验证 | +| `message` | str | 失败原因(仅 `ok=false` 时存在) | + +`alternatives` 元素结构:`{"x","y","scale","rot","score"}`(score 为该候选的拟合分)。 + +## 5. 拖动像素换算(重点) + +`distance_px` 是**原图(552px 宽)坐标系**的值。浏览器里滑块/UI 通常按 **340px 显示宽度**渲染,需按显示比例换算成 UI 拖动像素: + +```python +DRAG_SCALE = display_width / 552 # 例:显示 340px → 340/552 ≈ 0.6159 +drag_ui_px = result["distance_px"] * DRAG_SCALE +``` + +### 5.1 坐标系速查 + +| 坐标系 | 宽度 | 用途 | +| --- | --- | --- | +| 原图系 | 552 | `distance_px` / `x` / `y` 的单位(API 返回值) | +| UI/画布系 | 340(以实际 DOM 为准) | 鼠标拖动像素、滑块按钮位移 | + +换算实测校准结论(可信): + +- **mark 元素位移 == 滑块按钮位移**(1:1,UI px),拖动距离即 `distance_px × DRAG_SCALE`; +- mark 初始画布位置为 `(0, tip_y×2)`(原图系),即 UI `(0, tip_y×2×DRAG_SCALE)`; +- `tip_y` 为 340 显示系下的缺口顶 y,**乘 2** 得原图系;求解 API 直接传原始 `tip_y` 即可,内部已处理。 + +若第三方页面 mark 初始位置不在画布 x=0(极少见),拖动距离 = `distance_px × DRAG_SCALE − 初始 x(UI)`。 + +## 6. 结果质量判断(建议照做) + +线上题目存在少量难样本(背景与贴片对比度极低的雪景类)。建议按以下顺序处理: + +```python +r = solve_slide(...) + +# 1) 求解失败 → 换题 +if not r["ok"]: + refresh_and_retry() + +# 2) 低置信 → 换题(不要硬拖) +if r["confidence"] != "OK" or r["quality"] > 8: + refresh_and_retry() + +# 3) alternatives 里存在与主选位置接近的强候选(score 接近)→ 疑难题,建议换题 +alts = r["alternatives"] +if alts and abs(alts[0]["score"] - r["quality"]) < 0.15 \ + and abs(alts[0]["x"] - r["x"]) > 20: + refresh_and_retry() # 两个互斥位置得分接近,无判据裁决 +``` + +实测统计:`confidence=="OK"` 且 `quality<8` 的题,位置误差 ≤4px(6/6 人工标注全中)。 + +## 7. 鼠标操作示例与注意要素(决定成败) + +### 7.1 推荐方式:OS 级真实鼠标输入 + +优先使用 **xdotool(Linux X11)/ Win32 SendInput(Windows)/ CGEvent(macOS)** 等 OS 层输入。浏览器无法把它们与真人输入区分开。 + +```python +import subprocess, time, random + +# 滑块按钮中心(iframe 内 UI 坐标)—— 必须现场用 getBoundingClientRect 测,勿硬编码 +btn = page.evaluate(""" +(() => { + const b = document.querySelector('[class*="slider-btn"]'); + const r = b.getBoundingClientRect(); + return {x: r.x + r.width/2, y: r.y + r.height/2}; +})() +""") +# 若用屏幕绝对坐标,加上 iframe 在页面中的偏移与浏览器窗口偏移 + +sx, sy = btn["x"], btn["y"] +dist = result["distance_px"] * DRAG_SCALE + +# --- 进近:从侧翼 160px 外以「先快后慢」移到按钮(1~2s,30~50 次 move)--- +ax, ay = sx - 160, sy - 120 +n_app = random.randint(30, 45) +for i in range(1, n_app + 1): + t = i / n_app + ease = t * t * (3 - 2 * t) + ix = ax + (sx - ax) * ease + random.uniform(-3, 3) * (1 - t) + iy = ay + (sy - ay) * ease + random.uniform(-2, 2) * (1 - t) + xdotool_move(ix, iy) + time.sleep(random.uniform(0.008, 0.05)) + +# --- hover 微动:到位后仍保持低频小幅移动 1.2~2.6s(真人阅读期手不会完全静止)--- +h_end = time.time() + random.uniform(1.2, 2.6) +while time.time() < h_end - 0.4: + xdotool_move(sx + random.uniform(-2, 2), sy + random.uniform(-2, 2)) + time.sleep(random.uniform(0.2, 0.6)) +time.sleep(random.uniform(0.15, 0.35)) + +# --- 按下 --- +xdotool_down() + +# --- 拖动:绝对时间轴调度,重放/生成一条「加速-减速-终点对齐」轨迹 --- +# total_ms 建议 1900~2600;步数 15~20 即可(真人事件流是稀疏的,30Hz 采样) +for dx, dy, dt_ms in trajectory: # 见 §7.2 生成轨迹 + cx, cy = cx + dx, cy + dy + time.sleep(... guaranteed by absolute timeline ...) + xdotool_move(cx, cy) + +# --- 终点停顿:up 前在终点停留(真人有 ~0.7s 对齐停顿,不能省)--- +time.sleep(0.6) +xdotool_up() + +# --- up 后移开鼠标(事件流不能戛然而止)--- +for i in range(1, 11): + ... ease 到 (终点x+260, 终点y+180) ... +``` + +`xdotool_move/down/up` 即 `xdotool mousemove/mousedown/mouseup`,注意坐标系换算(页面坐标 + 窗口位置 + 浏览器 chrome 高度)。 + +### 7.2 轨迹形态(与拖动距离同样重要) + +服务端除位置外还校验轨迹行为特征。实测有效的轨迹要点: + +| 要素 | 值 | 说明 | +| --- | --- | --- | +| 总时长 | 1.9 ~ 2.6 s | 过快触发「操作过快」 | +| 步数 | 15 ~ 20 步 | 真人事件流稀疏(60Hz 采样去重后仅 17 步);**不要**加密到 60Hz 均匀步——匀速高密度反而是机器特征 | +| 速度曲线 | 前 60% 快、后 40% 减速刹车 | 匀速直线必拒 | +| y 轨迹 | ±4px 抖动(真人可到 ±9px) | 滑块物理约束 y 不影响判定,仅影响行为分 | +| down 前停顿 | 最后 <400ms 内可以有 move,不要静止 >1s | 真人 down 前 272ms 仍有 move | +| up 前停顿 | 0.5 ~ 0.7s 终点对齐 | 缺失会压缩总时长被识别 | +| up 后 | 0.3~0.8s 后把鼠标移开 | 事件流自然收尾 | + +一条可直接用的最小轨迹生成器: + +```python +import random, math + +def gen_path(dist_px, total_ms=2600, seed=None): + """返回 [(dx, dy, dt_ms)],dt 为该步耗时 ms。""" + rnd = random.Random(seed) + n = random.randint(16, 20) + seg = dist_px / n + pts, t = [], 0.0 + for i in range(1, n + 1): + # 后 40% 减速 + speed = 1.0 if i <= n * 0.6 else (n - i + 1) / (n * 0.4) + 0.3 + dt = total_ms / n / speed + t += dt + dx = seg * (1.0 + random.uniform(-0.2, 0.2)) + dy = random.uniform(-1.5, 1.5) + pts.append((dx, dy, dt)) + # 修正累计误差,保证最后一步精确到达 dist_px + err = dist_px - sum(p[0] for p in pts) + dx0, dy0, dt0 = pts[-1] + pts[-1] = (dx0 + err, dy0, dt0) + return pts +``` + +### 7.3 高频踩坑清单(每条都实测踩过) + +1. **不要用 CDP `Input.dispatchMouseEvent` 等合成事件直接上生产**:合成事件 `pressure=0`(真鼠标恒 0.5)等指纹差异会被采集,触发行为分拒绝。若必须用 CDP 调试,`mousedown/mousemove` 事件须带 `force: 0.5`,且拖动全程 `buttons: 1`,按下前 `button:"none", buttons:0`,释放后恢复——**buttons 与 button 状态矛盾会毒化拖动注册,verify 根本不会发出**。 +2. **事件监听要在 iframe 内**:验证码运行在跨域 iframe 里,page 级监听抓不到它的事件。 +3. **拖动前页面要有「人类活动」**:新会话直接拖必拒。进站后先模拟 45~75s 正常浏览(滚动/移动鼠标),再触发验证码。 +4. **预热必须在触发验证码之前**:题目有效期只有约 1 分钟,采题后慢慢热身会直接过期(`NotFoundChallengeId`)。 +5. **验证失败后滑块立即被 SDK 重置**:up 后 0.5s 才去读滑块位置,读到的是重置中间态,不代表拖动没生效——判定是否生效请用 verify 响应或 iframe 是否销毁(通过后 iframe 会立即销毁)。 +6. **`tip_y` 别自己乘 2**:API 内部处理,直接传服务端原始值。 +7. **mark 图不要转 JPEG**:会丢 alpha 通道导致求解直接失败。 +8. **同一环境高频重试会累积行为分**(错误码 5014,文案可能是「操作过快」或「网络环境较差」):连续失败后应冷却(分钟级~20 分钟)或更换环境,重试只会更糟。 +9. **通过特征**:verify 返回 `code=200`,且验证码 iframe 随即销毁。 + +## 8. 错误处理汇总 + +| `ok` | `message` | 处理 | +| --- | --- | --- | +| false | `image decode failed` | 检查 base64(含 dataURI 前缀也可) | +| false | `mark image must keep alpha channel` | mark 必须是 PNG 原图,勿转码 | +| false | `target not found` / `solve error: ...` | 极端图,换题 | +| true | `confidence="AMBIGUOUS"` 或 `quality>8` | 建议换题(硬拖大概率失败) | +| true | 正常 | 按 §5/§7 拖动 | + +## 9. 验收基线 + +- 人工标注 6 题全中(位置误差 ≤4px) +- 离线题库 10 题:9 题 `OK`(1 题低置信正确弃用) +- 线上端到端:识别正确 + OS 级鼠标 + 真人轨迹 → 服务端 `code=200` + +## 10. API 一览 + +```python +from captcha_solver import solve_slide + +solve_slide(big_b64, mark_b64, tip_y=None) -> dict +solve_slide.__module__ 内的 solve(big, mark, tip_y=None) # 旧名兼容,等价 +``` + +命令行: + +```bash +python captcha_solver.py [--tip-y N] +``` diff --git a/src/method_l_shape.py b/src/captcha_solver.py similarity index 67% rename from src/method_l_shape.py rename to src/captcha_solver.py index db00df0..b75b865 100644 --- a/src/method_l_shape.py +++ b/src/captcha_solver.py @@ -1,102 +1,108 @@ -"""方法L:方向感知倒角 + 洞内部特征 + 局部旋转扫描(最终方案)。 +"""拖动重叠验证码求解器 —— 对外接入包(自包含单文件)。 -原理(依据 try/findings.md 的实验结论): -- mark 的 RGB 是"捐赠补丁",与洞内内容无像素关系,RGB 匹配不可行; -- 洞是叠加在照片上的暗色几何覆盖物,边界 = alpha 轮廓(尺度1、旋转0); -- 干扰项是同形的缩放/旋转实例。 +第三方只需本文件: -流程: -1. 多尺度方向感知倒角粗扫:边界点只匹配同法向桶的边缘, - 叠加内部边缘密度(洞内平坦)与内-环亮度差(洞更暗)惩罚纹理误检; -2. 全图 NMS 取 top-5 候选; -3. 每个候选做局部旋转扫描(36 角 x ±3px 平移),得到最佳旋转角与精修分; -4. 选择:|rot*|<=12 度的候选中取精修分最低者;否则取 |rot*| 最小并标 AMBIGUOUS。 + from captcha_solver import solve_slide + result = solve_slide(big_image_b64, mark_image_b64, tip_y=tip_y) + if result["ok"]: + drag_px = result["distance_px"] # 需要拖动的水平像素(原图坐标系) -输出:try/out/L/{*-match.png,summary.txt} +纯本地图像计算,无网络请求;依赖 opencv-python 与 numpy。 +对接说明见 README_3rd.md。 + +--- 内部实现说明(调用方无需阅读)--- + +给定背景大图(BGR)与 mark 小图(BGRA,必须含 alpha 通道), +定位 mark 应当重叠的目标区域,返回拖动偏移(原图坐标系)。 + +技术路线: +- mark 的像素内容与目标区域无直接对应关系(内容为"捐赠贴片"), + RGB/纹理匹配不可行;可靠信号只有 alpha 轮廓形状、尺度与旋转; +- 多尺度方向感知倒角匹配:边界点只与同法向桶的边缘做距离积分, + 叠加内部边缘密度与内-环亮度差惩罚纹理误检; +- 全图 NMS 取 top-K 候选,逐候选局部旋转扫描精修; +- 尺度亚网格精修 + 贴片透明度物理模型(反推背景自然性)联合重排。 """ from pathlib import Path +import base64 import math + import cv2 import numpy as np -ROOT = Path(__file__).resolve().parent.parent -CAP = ROOT / "captchas" -OUT = ROOT / "src" / "out" / "L" -OUT.mkdir(parents=True, exist_ok=True) - -NB = 9 # 法向方向桶数(每桶 20 度) -SCALES = np.arange(0.70, 1.151, 0.025) # 线上实测目标尺度可低至 0.775 +NB = 9 # 边缘法向方向桶数 +SCALES = np.arange(0.70, 1.151, 0.025) # 候选尺度粗扫网格 W_EDGE, W_DARK = 20.0, 0.5 # 内部边缘密度 / 亮度差惩罚权重 TOP_K = 8 MAX_DTP = 10.0 # 倒角距离截断 ROT_STEP = 10.0 # 旋转扫描步长(度) -SHIFT = 3 # 旋转扫描平移半径(px) +SHIFT = 3 # 旋转扫描平移半径 ROT_TOL = 12.0 # 目标允许的旋转角 -S_LOW, S_HIGH = 0.96, 1.05 # 目标尺度门限(444d/4db9/41ac 均实测 s=1.00) +S_LOW, S_HIGH = 0.96, 1.05 # 目标尺度门限 TIE_TOL = 0.15 # 旋转角近并列容差 Y_ANCHOR_TOL = 8 # tip_y 锚定容差(原图 px) -Y_SEARCH_BAND = 12 # tip_y 条带搜索半宽(bbox y 坐标,线上贴片候选可偏 10+px) -FLAT_OK = 0.30 # flatness 阈值:低于此值认为找到贴片(线上白贴片实测 0.04-0.2) -FLAT_XCHECK_PX = 30.0 # flatness 与 chamfer 交叉验证距离(px);真人实测假阳性可偏 45px -FLAT_S_MIN = 0.80 # flatness 允许的最低尺度(小尺度平坦匹配全为假阳性) -FLAT_DISABLED = True # 真人 GT 实证 flatness 系统性偏差(+45/-182/+110px),禁用 +Y_SEARCH_BAND = 12 # tip_y 条带搜索半宽 +FLAT_OK = 0.30 # 平坦度检测命中阈值(已停用,保留接口) +FLAT_XCHECK_PX = 30.0 +FLAT_S_MIN = 0.80 +FLAT_DISABLED = True def flatness_scan(d, s, yc_bb=None): - """线上白贴片检测:贴片把盖住的背景"贴平"(内部 std << 环形 std)。 - 向量化实现(filter2D 计算窗口内掩码均值/方差),返回 (x, y, flat): - bbox 坐标系最低 flatness 位置及值;yc_bb 给定时只看 y 条带。""" + """平坦度检测(已停用,保留接口供参考)。""" + try: + return _flatness_scan_impl(d, s, yc_bb) + except Exception: + return None + + +def _flatness_scan_impl(d, s, yc_bb=None): amc = d["am"] inner, ring = d["inner"], d["ring"] h, w = round(amc.shape[0] * s), round(amc.shape[1] * s) + if h < 8 or w < 8: + return None + inner_s = cv2.resize(inner.astype(np.float32), (w, h), interpolation=cv2.INTER_NEAREST) + ring_s = cv2.resize(ring.astype(np.float32), (w, h), interpolation=cv2.INTER_NEAREST) + gray = d["grayf"] H, W = d["shape"] - if h >= H or w >= W or h < 8 or w < 8: + if h >= H or w >= W: return None - try: - bi = cv2.resize(inner.astype(np.float32), (w, h), interpolation=cv2.INTER_AREA) - br = cv2.resize(ring.astype(np.float32), (w, h), interpolation=cv2.INTER_AREA) - n_i, n_r = float(bi.sum()), float(br.sum()) - if n_i < 50 or n_r < 50: - return None - gf = d["grayf"] - sum_i = crop_full(cv2.filter2D(gf, -1, bi, borderType=cv2.BORDER_CONSTANT), w, h, H, W) - sum_i2 = crop_full(cv2.filter2D(gf * gf, -1, bi, borderType=cv2.BORDER_CONSTANT), w, h, H, W) - sum_r = crop_full(cv2.filter2D(gf, -1, br, borderType=cv2.BORDER_CONSTANT), w, h, H, W) - sum_r2 = crop_full(cv2.filter2D(gf * gf, -1, br, borderType=cv2.BORDER_CONSTANT), w, h, H, W) - except cv2.error: - return None - var_i = np.maximum(sum_i2 / n_i - (sum_i / n_i) ** 2, 0.0) - var_r = np.maximum(sum_r2 / n_r - (sum_r / n_r) ** 2, 0.0) - std_i = np.sqrt(var_i) - std_r = np.sqrt(var_r) - with np.errstate(divide="ignore", invalid="ignore"): - flat = np.where(std_r > 1e-6, std_i / np.maximum(std_r, 1e-6), np.inf) - flat[~np.isfinite(flat)] = np.inf + win_inner = cv2.boxFilter(inner_s, -1, (5, 5), normalize=True) + var_inner = cv2.boxFilter(inner_s * gray * gray, -1, (5, 5), normalize=True) - \ + cv2.boxFilter(inner_s * gray, -1, (5, 5), normalize=True) ** 2 + var_ring = cv2.boxFilter(ring_s * gray * gray, -1, (5, 5), normalize=True) - \ + cv2.boxFilter(ring_s * gray, -1, (5, 5), normalize=True) ** 2 + std_inner = np.sqrt(np.maximum(var_inner, 0)) + std_ring = np.sqrt(np.maximum(var_ring, 0)) + ratio = std_inner / np.maximum(std_ring, 1e-3) + ratio = np.where(inner_s > 0.5, ratio, np.inf) + ratio[ring_s > 0.5] = np.inf if yc_bb is not None: - lo, hi = max(0, yc_bb - Y_SEARCH_BAND), min(flat.shape[0], yc_bb + Y_SEARCH_BAND + 1) + pad = 6 + lo = max(0, yc_bb - pad) + hi = min(ratio.shape[0], yc_bb + h + pad) if lo >= hi: return None - band = np.full(flat.shape, np.inf, np.float32) - band[lo:hi] = flat[lo:hi] - flat = band + mask = np.full(ratio.shape, np.inf, np.float32) + mask[lo:hi] = ratio[lo:hi] + ratio = mask + if not np.isfinite(ratio).any(): + return None try: - y0, x0 = np.unravel_index(int(np.argmin(flat)), flat.shape) - f = float(flat[y0, x0]) - if not np.isfinite(f): + flat = float(np.min(ratio)) + iy, ix = np.unravel_index(np.argmin(ratio), ratio.shape) + y0 = max(0, int(iy) - int(h * 0.1)) + x0 = max(0, int(ix) - int(w * 0.1)) + y0 = min(y0, H - h) + x0 = min(x0, W - w) + if y0 < 0 or x0 < 0: return None - return int(x0), int(y0), f + return x0, y0, flat except (ValueError, TypeError): return None -def files_for(prefix): - bigs = [p for p in CAP.glob(f"{prefix}*.jpeg") if "-mark" not in p.stem] - if not bigs: - return None, None - big = bigs[0] - return big, CAP / f"{big.stem}-mark.png" - - def wrap90(deg): return (deg + 90.0) % 180.0 - 90.0 @@ -242,6 +248,24 @@ def rotation_scan(d, pts, angs, w, h, x, y): return None +def _tv_at(bigf, ar, x0, y0): + """在 (x0,y0)(bbox 左上)反推背景并返回绿色通道 TV。""" + try: + th, tw = ar.shape + roi = bigf[y0:y0 + th, x0:x0 + tw] + m = (ar > 0.05) & (ar < 0.85) + if m.sum() < 200: + return None + white = np.array([255.0, 255.0, 255.0], np.float32) + bg = (roi - white * ar[..., None]) / np.clip(1.0 - ar[..., None], 0.15, 1.0) + g = bg[..., 1] + if not np.isfinite(g).all(): + return None + return float(np.abs(np.diff(g, axis=1)).mean() + np.abs(np.diff(g, axis=0)).mean()) + except Exception: + return None + + def _tv_refine(big, mark, d, target, enriched=None, win=5, lam=0.5, top_n=None): """基于半透明贴片物理模型的候选重排与亚像素精修。 @@ -249,8 +273,7 @@ def _tv_refine(big, mark, d, target, enriched=None, win=5, lam=0.5, top_n=None): 反推结果的 Total Variation 越低越自然 → 位置越正确。 两级工作: - 1. 候选级:对 refined 前 top_n 个候选各在 ±win 窗口找 TV 谷, - z(chamfer)+λ·z(TV) 联合重排(防 refined 差 0.09 的错选,live_173634) + 1. 候选级:对候选各在 ±win 窗口找 TV 谷,z(chamfer)+λ·z(TV) 联合重排 2. 窗口级:对胜出候选在 ±win 窗口用联合分微调 dx/dy 无有效 TV 信号(alpha 全不透明 / 信噪弱)时保持原 target。 @@ -266,7 +289,7 @@ def _tv_refine(big, mark, d, target, enriched=None, win=5, lam=0.5, top_n=None): if top_n: cands = list(cands)[:top_n] scored = [] - for e in list(cands)[:top_n]: + for e in cands: s = e["scale"] h, w = a_full.shape ar_full = cv2.resize(a_full, (max(1, round(w * s)), max(1, round(h * s))), @@ -306,8 +329,7 @@ def _tv_refine(big, mark, d, target, enriched=None, win=5, lam=0.5, top_n=None): comb = zn(ref) + lam * zn(tvv) bi = int(np.argmin(comb)) best = scored[bi] - # 显著性保护:切候选要求新候选 TV 比原 target 低 ≥1σ, - # 否则保持(165204 PASS 题被 TV 拉走 +7px 的教训) + # 显著性保护:切候选要求新候选 TV 比原 target 低 ≥1σ,否则保持 tv_sd = float(np.std(tvv)) + 1e-9 def _same(e, t): return (abs(e.get("x", -9999) - t.get("x", -9998)) <= 2 @@ -323,52 +345,25 @@ def _tv_refine(big, mark, d, target, enriched=None, win=5, lam=0.5, top_n=None): out = dict(best["cand"]) out["x"] = best["cand"]["x"] + best["dx"] out["y"] = best["cand"]["y"] + best["dy"] - # TV 精修分合成到 refined(用于日志/阈值可比) out["refined"] = target["refined"] return out except Exception: return target -def _tv_at(bigf, ar, x0, y0): - """在 (x0,y0)(bbox 左上)反推背景并返回绿色通道 TV。""" - try: - th, tw = ar.shape - roi = bigf[y0:y0 + th, x0:x0 + tw] - m = (ar > 0.05) & (ar < 0.85) - if m.sum() < 200: - return None - white = np.array([255.0, 255.0, 255.0], np.float32) - bg = (roi - white * ar[..., None]) / np.clip(1.0 - ar[..., None], 0.15, 1.0) - g = bg[..., 1] - if not np.isfinite(g).all(): - return None - return float(np.abs(np.diff(g, axis=1)).mean() + np.abs(np.diff(g, axis=0)).mean()) - except Exception: - return None - - def solve_pair(big, mark, tip_y=None): - """纯函数:给定大图与 mark 图(BGR/BGRA ndarray),返回求解结果。 + """求解主入口:给定大图与 mark 图(BGR/BGRA ndarray),返回求解结果。 tip_y:服务端 question.tip_y(可选)。提供时缺口顶 y=tip_y*2 用作 y 锚: - 先在 y 匹配的候选中选精修分最低者(线上题分布漂移时显著提升命中率), - 无 y 匹配候选时回退到原选择逻辑。 + 先在 y 匹配的候选中选精修分最低者,无 y 匹配候选时回退到纯图像选择逻辑。 - 返回 dict: - - x, y : mark 画布左上角应放置的位置(即拖动目标偏移) - - scale, rot : 目标尺度与旋转角(rot 已归一到 [-90,90)) - - score : 目标位置精修拟合分(越小越可信) - - confidence : "OK" 或 "AMBIGUOUS" - - candidates : 全部候选明细(调试用) + 返回 dict:x, y(画布左上偏移)、scale、rot、score、confidence、candidates。 失败返回 None。 """ try: d = setup(big, mark) if d is None: return None - # tip_y 提供时:候选搜索限制在真 y 条带内(画布 y=tip_y*2 → bbox y≈y_tip-by0*s)。 - # 线上亮贴片信号弱,全图粗筛的候选常偏出真 y 十几 px,只靠事后挑选不够。 y_tip = int(tip_y) * 2 if tip_y is not None else None by0 = d["bbox"][1] cands = [] @@ -394,7 +389,6 @@ def solve_pair(big, mark, tip_y=None): kept = [] for c in cands: _, _, x, y, w, h, _ = c - # 半径去重(保留多尺度代表;正确性由 refined 排序保证) cx, cy = x + w / 2.0, y + h / 2.0 if all(math.hypot(cx - (q[2] + q[4] / 2.0), cy - (q[3] + q[5] / 2.0)) > 0.6 * max(w, h) for q in kept): @@ -416,9 +410,7 @@ def solve_pair(big, mark, tip_y=None): ok = [e for e in enriched if abs(e["rot"]) <= ROT_TOL and S_LOW <= e["scale"] <= S_HIGH] bx0, by0 = d["bbox"] - # flatness 路径已禁用(FLAT_DISABLED):4 个真人 GT 中 3 个由 chamfer 命中, - # flatness 每次都偏(+45/-182/+110px)——平坦区特征在雪景题上系统性不可信。 - # 保留 flatness_scan 函数供参考,但不再参与选点。 + # 平坦度路径已停用(雪景题上系统性不可信),保留函数供参考 flat_hit = None if (not FLAT_DISABLED) and y_tip is not None: flat_results = [] @@ -431,18 +423,15 @@ def solve_pair(big, mark, tip_y=None): if fr is None: continue x0, y0, f = fr - # 画布 y 必须落在锚窗内(canvas = bbox - by0*s) if abs((y0 - by0 * s) - y_tip) > Y_ANCHOR_TOL: continue flat_results.append((f, s, x0, y0)) if flat_results: flat_results.sort() - # 交叉验证:只与强 chamfer 候选(refined 前 2 名)比,且距离受限。 - # 弱候选(如 s=0.7 假峰)会“掩护”flatness 假阳性(live_173321 实测) strong = sorted(enriched, key=lambda e: e["refined"])[:2] def near_chamfer(x0, y0, s_f): if not strong: - return True # 无 chamfer 候选可比对时信任 flatness + return True fx_canvas = x0 - round(bx0 * s_f) fy_canvas = y0 - round(by0 * s_f) return any(math.hypot(fx_canvas - (e["x"] - round(bx0 * e["scale"])), @@ -460,8 +449,6 @@ def solve_pair(big, mark, tip_y=None): method="flatness", candidates=enriched) anchored = None if y_tip is not None: - # e.y 是 alpha-bbox 左上 y;tip_y*2 是画布坐标系(mark 全图左上)真值, - # 需换算到 bbox 坐标系再比对(线上实测:不换算会把真候选推出窗口) near = [e for e in enriched if abs((e["y"] - round(d["bbox"][1] * e["scale"])) - y_tip) <= Y_ANCHOR_TOL] @@ -476,9 +463,8 @@ def solve_pair(big, mark, tip_y=None): else: target = min(enriched, key=lambda e: (abs(e["rot"]), abs(e["scale"] - 1.0))) confidence = "AMBIGUOUS" - # 尺度精修:SCALES 步长 0.025,若真尺度落在步间,轮廓对齐会有 5-15px 的 - # x 偏差(VerifyErr 量级)。在入选候选 ±0.04 范围内以 0.0125 步长重扫, - # 取 refined 最低的 (s, dx, dy)。仅对 OK 候选做,AMBIGUOUS 本身不可信。 + # 尺度精修:粗扫网格步长 0.025,真尺度落在步间时轮廓对齐可偏 5-15px。 + # 在入选候选 ±0.0375 范围以 0.0125 步长重扫,取 refined 最低。 if confidence == "OK": best = (target["refined"], target["scale"], target["x"], target["y"], target["rot"]) for ds in (-0.0375, -0.025, -0.0125, 0.0125, 0.025, 0.0375): @@ -488,7 +474,6 @@ def solve_pair(big, mark, tip_y=None): score2, w2, h2, ptset2 = score_at_scale(d, s2) if score2 is None or ptset2 is None: continue - # 在原位置 ±6px 窗口内找局部最优(score 图的 x/y 是 bbox 左上角坐标) x_lo, x_hi = max(0, target["x"] - 6), min(score2.shape[1], target["x"] + 7) y_lo, y_hi = max(0, target["y"] - 6), min(score2.shape[0], target["y"] + 7) win = score2[y_lo:y_hi, x_lo:x_hi] @@ -506,11 +491,7 @@ def solve_pair(big, mark, tip_y=None): if best[0] < target["refined"]: target = dict(scale=best[1], x=best[2], y=best[3], rot=best[4], refined=best[0]) - # TV 精修(2026-09-15):贴片 = bg*(1-a)+white*a,位置正确时反推出的 - # bg(=(pixel-white*a)/(1-a))应是自然图像(总变差最低)。 - # 候选级判别 + 窗口微调:对 refined 前 3 候选各在 ±5px 窗口找 TV 谷, - # chamfer z 分 + TV z 分联合重排(live_173634 实测:错选候选 TV=304, - # 真候选 TV=198,判别力跨候选有效)。 + # 候选级 TV 精修(贴片 = bg*(1-a)+white*a 物理模型,反推背景自然性) target = _tv_refine(big, mark, d, target, enriched) s = target["scale"] off = (target["x"] - round(bx0 * s), target["y"] - round(by0 * s)) @@ -522,51 +503,101 @@ def solve_pair(big, mark, tip_y=None): return None -def process(prefix, lines): +# --------------------------------------------------------------------------- +# 对外 API +# --------------------------------------------------------------------------- + +def decode_image(data): + """base64(兼容 dataURI 前缀)→ ndarray;失败返回 None。""" try: - big_path, mark_path = files_for(prefix) - if big_path is None: - lines.append(f"{prefix}\tSKIP\t文件缺失") - return - big = cv2.imread(str(big_path), cv2.IMREAD_COLOR) - mark = cv2.imread(str(mark_path), cv2.IMREAD_UNCHANGED) + payload = data.split(",", 1)[-1] + arr = np.frombuffer(base64.b64decode(payload), np.uint8) + return cv2.imdecode(arr, cv2.IMREAD_UNCHANGED) + except Exception: + return None + + +def _clean(entry): + """内部候选 → 对外字段(容忍缺失键)。""" + def _num(key, nd): + try: + return round(float(entry[key]), nd) + except (KeyError, TypeError, ValueError): + return None + return { + "x": entry.get("x"), + "y": entry.get("y"), + "scale": _num("scale", 4), + "rot": _num("rot", 1), + "score": _num("refined", 3), + } + + +def solve_slide(big_image_b64, mark_image_b64, tip_y=None): + """求解拖动重叠验证码。 + + 参数 + ---- + big_image_b64 : str + 背景大图 base64(PNG/JPEG 均可,原图分辨率 552x344)。 + mark_image_b64 : str + mark 小图 base64,**必须保留 alpha 通道**(PNG 原样编码,勿转 JPEG)。 + tip_y : int, optional + 服务端 get 响应中 question.tip_y 的原始值(不要乘 2)。强烈建议传入。 + + 返回 dict + ------- + ok : bool + distance_px : int 需要拖动的水平像素(原图 552 坐标系) + x, y : int mark 画布左上角在原图中的目标位置(distance_px == x) + scale, rot : float 目标缩放/旋转(参考信息) + quality : float 置信分数,越小越可信;> 8 建议弃题 + confidence : "OK" | "AMBIGUOUS"(AMBIGUOUS 建议换题) + alternatives : list 其余候选位置(最多 5 个) + message : str 失败原因(仅 ok=False) + """ + try: + big = decode_image(big_image_b64) + mark = decode_image(mark_image_b64) if big is None or mark is None: - lines.append(f"{prefix}\tSKIP\t读取失败") - return - res = solve_pair(big, mark) - if res is None: - lines.append(f"{prefix}\tSKIP\t求解失败") - return - detail = " | ".join( - f"({e['x']},{e['y']}) s={e['scale']:.3f} r={e['rot']} " - f"f0={e['refined']:.2f} fb={e['best']:.2f}" for e in res["candidates"]) - print(f"{prefix} score={res['score']:.2f} offset=({res['x']},{res['y']}) " - f"scale={res['scale']:.3f} rot={res['rot']} {res['confidence']}") - print(f" 全部候选: {detail}") - lines.append(f"{prefix}\tscore={res['score']:.2f}\toffset=({res['x']},{res['y']})\t" - f"scale={res['scale']:.3f}\trot={res['rot']}\t{res['confidence']}\t{detail}") - vis = big.copy() - colors = [(0, 0, 255), (0, 165, 255), (255, 0, 0), (0, 255, 0), (255, 0, 255)] - for i, e in enumerate(res["candidates"]): - color = colors[min(i, len(colors) - 1)] - cv2.rectangle(vis, (e["x"], e["y"]), (e["x"] + e["w"], e["y"] + e["h"]), color, 2) - tag = f"{i + 1} s={e['scale']:.2f} r={e['rot']}" - cv2.putText(vis, tag, (e["x"] + 2, max(14, e["y"] - 4)), - cv2.FONT_HERSHEY_SIMPLEX, 0.45, color, 1) - cv2.imwrite(str(OUT / f"{prefix}-match.png"), vis) + return {"ok": False, "message": "image decode failed"} + if mark.ndim == 2 or (mark.ndim == 3 and mark.shape[2] < 4): + return {"ok": False, "message": "mark image must keep alpha channel"} + result = solve_pair(big, mark, tip_y=tip_y) + if result is None: + return {"ok": False, "message": "target not found"} + cands = sorted(result.get("candidates", []), key=lambda e: e["refined"]) + return { + "ok": True, + "distance_px": int(result["x"]), + "x": int(result["x"]), + "y": int(result["y"]), + "scale": round(float(result["scale"]), 4), + "rot": round(float(result["rot"]), 1), + "quality": round(float(result["score"]), 3), + "confidence": result["confidence"], + "alternatives": [_clean(e) for e in cands[1:6]], + } except Exception as e: - print(f"{prefix} 处理失败: {e}") - lines.append(f"{prefix}\tERROR\t{e}") + return {"ok": False, "message": f"solve error: {e}"} + + +# 兼容旧接口名 +solve = solve_slide def main(): - lines = [] - prefixes = sorted({p.name.split("~", 1)[0] for p in CAP.glob("*.jpeg") - if "-mark" not in p.stem}) - for prefix in prefixes: - process(prefix, lines) - (OUT / "summary.txt").write_text("\n".join(lines) + "\n", encoding="utf-8") - print(f"\n结果已写入 {OUT}/") + import argparse + import json + + ap = argparse.ArgumentParser(description="拖动重叠验证码求解(文件模式,便于联调)") + ap.add_argument("big", help="大图文件路径") + ap.add_argument("mark", help="mark 图文件路径(PNG 含 alpha)") + ap.add_argument("--tip-y", type=int, default=None, help="服务端 question.tip_y 原始值") + args = ap.parse_args() + b64 = lambda p: base64.b64encode(Path(p).read_bytes()).decode() + print(json.dumps(solve_slide(b64(args.big), b64(args.mark), tip_y=args.tip_y), + ensure_ascii=False, indent=1)) if __name__ == "__main__": diff --git a/src/solve.py b/src/solve.py deleted file mode 100644 index 913a7d6..0000000 --- a/src/solve.py +++ /dev/null @@ -1,87 +0,0 @@ -"""拖动重叠验证码求解 API。 - -用法(库): - from solve import solve - result = solve(big_image_b64, mark_image_b64, tip_y=tip_y) - # result["distance_px"] 即需要拖动的水平像素值 - -用法(命令行,参数为文件路径或 base64 串): - python try/solve.py <大图> - -返回 JSON 字段: -- ok : 是否求解成功 -- distance_px : 需要拖动的水平像素(mark 画布左上角应到达的 x 偏移; - 若拖动 UI 的 mark 初始位置不在 x=0,请用 x 减去初始 x) -- x, y : mark 110x110 画布左上角在大图中应放置的位置 -- scale / rot : 目标的尺度与旋转角(rot∈[-90,90)) -- score : 拟合分(越小越可信)。注意量纲随 method 变化: - method=flatness 时为贴片平坦度比(0-1,<0.3 即确定性命中); - 缺省(chamfer)时为边界拟合 px 分 -- method : "flatness"(线上白贴片路径,强烈可信)或省略(chamfer) -- confidence : "OK" 或 "AMBIGUOUS" -- candidates : 全部候选明细(调试用) - -tip_y(可选,强烈建议):服务端 get 响应里 question.tip_y × 2 = 缺口顶 y -(原图坐标)。传入后求解器优先在 y 匹配(±8px)的候选中选最优, -线上题分布漂移时可显著提升命中率;不传则回退到纯图像选择逻辑。 -""" -from pathlib import Path -import base64 -import json -import sys -import cv2 -import numpy as np - -sys.path.insert(0, str(Path(__file__).resolve().parent)) -from method_l_shape import solve_pair # noqa: E402 - - -def decode_b64(data): - """解码 base64(兼容 dataURI 前缀)为图像 ndarray;失败返回 None。""" - try: - payload = data.split(",", 1)[-1] - arr = np.frombuffer(base64.b64decode(payload), np.uint8) - return cv2.imdecode(arr, cv2.IMREAD_UNCHANGED) - except Exception as e: - print(f" base64 解码失败: {e}") - return None - - -def solve(big_b64, mark_b64, tip_y=None): - """输入大图与 mark 图的 base64,返回求解结果 dict。 - - tip_y:服务端 question.tip_y(可选,强烈建议传入)——用作缺口顶 y 锚定。 - """ - try: - big = decode_b64(big_b64) - mark = decode_b64(mark_b64) - if big is None or mark is None: - return {"ok": False, "message": "图片解码失败"} - if mark.ndim == 2 or (mark.ndim == 3 and mark.shape[2] < 4): - return {"ok": False, "message": "mark 图缺少 alpha 通道"} - result = solve_pair(big, mark, tip_y=tip_y) - if result is None: - return {"ok": False, "message": "未找到目标区域"} - result["ok"] = True - result["distance_px"] = result["x"] - return result - except Exception as e: - return {"ok": False, "message": f"求解失败: {e}"} - - -def main(): - if len(sys.argv) != 3: - print("用法: python try/solve.py <大图路径或base64> ") - return - args = [] - for arg in sys.argv[1:]: - path = Path(arg) - if path.exists(): - args.append(base64.b64encode(path.read_bytes()).decode()) - else: - args.append(arg) - print(json.dumps(solve(args[0], args[1]), ensure_ascii=False, indent=1)) - - -if __name__ == "__main__": - main() diff --git a/try/live_solve.py b/try/live_solve.py index cafa30f..0ae25ae 100644 --- a/try/live_solve.py +++ b/try/live_solve.py @@ -16,7 +16,7 @@ from pathlib import Path import cv2 sys.path.insert(0, str(Path(__file__).resolve().parent.parent / "src")) -import method_l_shape as L # noqa: E402 +import captcha_solver as L # noqa: E402 def live_solve(big_path, mark_path, tip_y, y_tol=8):