feat(任务): 录制手势改成**纯录制回放**(真手指轨迹)+ 手机端 getevent 录制

用户反馈:录制出来的滑动又被套上滑动那套逻辑(方向/幅度/拟人重新生成),
"导致滑动还是不顺畅"。录制就该是录制——按录下的路径与时间原样重放。

- 新步骤 `gesture`「录制手势」:存完整轨迹点列 `[[x,y,t_ms],…]`,
  回放时原样交给设备(`d.swipe_points(points, duration)`),不做任何加工。
  与 `swipe` 是两套东西(滑动是参数化的,录制是点列)。
- core/gesture.py(新):
  · 手机端录制 `PhoneRecorder`:`getevent` 读**真触屏**设备(自动挑 fts_ts 这类、
    排除 uinput/vitural-sar 等合成节点),解析两套协议(BTN_TOUCH / ABS_MT_TRACKING_ID);
    **合成的注入事件不会出现在真触屏节点上**(实测),所以录到的只有人手的动作;
  · 点列清洗:按 ≥16ms 抽稀但**末点必留**(快划时不能把收尾丢了);
  · 回放**一次调用**而不是逐点注入:实测这台设备单次触摸 RPC ≈190ms,
    逐点回放 20 点要 3.8 秒——只能让设备自己插值,"时间"由点密度还原。
- 接口:POST /api/gesture/record/{start,stop}(需设备权限;重复开始返回 409)。
- 「录制手势」步骤卡片:录到就回填(手机上录 / 网页上录),显示点数/时长;
  **滑动步骤上的「录制手势」按钮已移除**(按用户要求,录制不再走滑动逻辑)。
- 录制前自动唤醒设备:**息屏时 u2 抓 UI 树会从 3 秒退化到 65 秒**(实测),
  截图也是黑的——这是本次排查最耗时的坑,已写进文档。

自测:假设备单测(点列/时长/抽稀/失败退回直线,10 项);
无头浏览器 + CDP 真模拟拖动 → 录到 14 点/605ms 并按实际轨迹回填;
手机录制接口起停/重复拦截;
真机回放:用录下的轨迹跑任务 → 步骤 `gesture ok`。
⚠️ 手机端"真手指轨迹"这一段需要人真的去划才能端到端验证(我没有手指)。
文档:TASK_DEV §3.2(两套东西对比 + 三个坑)、步骤表 21 种、API、README。
This commit is contained in:
2026-09-20 15:04:21 +08:00
parent f62bc8da28
commit 73895f7233
11 changed files with 569 additions and 76 deletions
+276
View File
@@ -0,0 +1,276 @@
"""录制手势:**纯录制、纯回放**——录下真实手指的轨迹与时间,回放时一模一样地重放。
和「滑动」步骤的区别(这是两套东西,别混):
| | `swipe` 步骤 | `gesture` 步骤(本模块) |
|---|---|---|
| 内容 | 方向 + 幅度 + 时长区间 | **完整轨迹点列** `[[x, y, t_ms], ...]` |
| 回放 | 每次重新生成(弧线/抖动/本设备手速) | **照录制的路径与时间逐点重放**,不加任何修饰 |
| 适合 | 通用滑动,换设备也能用 | 你亲手划的、要求一模一样的手势 |
两个录制来源:
1. **手机上录**(`PhoneRecorder`):`getevent` 读**真触屏**设备,抓的是你手指的真实轨迹
——注意合成的注入事件不会出现在真触屏节点上(实测确认),所以录到的只有人手的动作;
2. **网页上录**:看着设备画面拖鼠标,前端采样路径与时间(见 `static/admin/recorder.js`)。
## 回放怎么做(实测定的,别改回去)
**一次调用把整条轨迹交给设备**(`d.swipe_points(points, duration)`),不是逐点注入。
原因:实测这台设备上**单次触摸注入 RPC ≈190ms**(`d.touch.move` 25 次平均 190ms),
逐点回放 20 个点要 3.8 秒——比录的手势慢一个数量级,根本不能用。所以只能让设备
自己去插值:`swipe_points` 内部按 `steps = duration/0.005` 在轨迹上均匀推进,
手势的真实速度由设备端决定。
拿到"时间"靠的是**点密度**:手指慢的地方采样点天然更密(单位长度上点更多),
无论设备是按弧长均匀插值还是按段均匀插值,"点密的段落走得更久"都成立
——所以把**录制的原始点 + 录制总时长**一起交给设备,速度曲线就能大致还原。
"""
import re
import subprocess
import threading
import time
from core.adb_helper import ADB_PATH
from core.logger import get_logger
_log = get_logger("core.gesture")
# 采样下限:手指在 16ms 内位移极小,比这更密的点只增加注入次数与延迟误差
MIN_GAP_MS = 16
MAX_POINTS = 240 # 单条手势的点数上限(约 4s @60Hz,够用且不让步骤 JSON 膨胀)
MAX_REPLAY_POINTS = 40 # 回放时最多发给设备多少个点(实测每点开销大,40 点是画质/耗时折中)
MIN_POINTS = 2
# getevent -lt 的行: [ 12345.678901] EV_ABS ABS_MT_POSITION_X 000003e8
# 第 4 列不能限定十六进制:BTN_TOUCH 的值是 DOWN/UP 这类**键名**(含非十六进制字母),
# 限定 [0-9a-fA-F]+ 会让这些行整个被跳过 → 老协议解析不出任何手势
_LINE_RE = re.compile(r"\[\s*([\d.]+)\]\s+(\S+)\s+(\S+)\s+(\S+)")
# "像真触屏"的设备名特征(先挑这些);明显是合成/虚拟的一律排除
_TOUCH_HINTS = ("ts", "touch", "goodix", "fts", "synaptics", "novatek", "himax", "sec_touch")
_VIRTUAL_HINTS = ("uinput", "virtual", "vitural", "sar", "gpio", "pon", "jack", "button")
# ================== 点列处理 ==================
def clean_points(points, min_gap_ms=MIN_GAP_MS, max_points=MAX_POINTS):
"""规范点列:丢掉过密的采样、限制总数,时间归零到第一点。
输入/输出都是 `[[x, y, t_ms], ...]`(t 为相对第一点的毫秒)。点数不足 2 返回 []。
"""
pts = []
for p in points or []:
try:
pts.append([int(p[0]), int(p[1]), float(p[2])])
except (TypeError, ValueError, IndexError):
continue
if len(pts) < MIN_POINTS:
return []
out = [pts[0]]
for p in pts[1:-1]:
if (p[2] - out[-1][2]) >= min_gap_ms:
out.append(p)
# 末点**必须保留**:它决定手势收尾在哪(快划时中间点稀,终点不能跟着丢)
if pts[-1][:2] != out[-1][:2] or pts[-1][2] > out[-1][2]:
out.append(pts[-1])
if len(out) < MIN_POINTS:
return []
if len(out) > max_points: # 均匀抽稀,保住首尾
step = len(out) / float(max_points)
out = [out[min(len(out) - 1, int(i * step))] for i in range(max_points)]
t0 = out[0][2]
for p in out:
p[2] = int(round(p[2] - t0))
return out
def duration_ms(points):
return int(points[-1][2]) if points else 0
def replay(d, points, speed=1.0):
"""按录制的轨迹与时长回放(一次调用交给设备插值)。返回发给设备的参数。
点太多会让设备端逐点开销累积(实测每个点约 1s 量级),所以超过 MAX_REPLAY_POINTS
就抽稀——路径形状基本不变,但快得多。首尾点一定保留。
"""
pts = clean_points(points, min_gap_ms=0) # 回放不再按时间丢点
if len(pts) < MIN_POINTS:
raise ValueError("手势点列无效(少于 2 个点)")
if len(pts) > MAX_REPLAY_POINTS:
n = MAX_REPLAY_POINTS
step = (len(pts) - 1) / float(n - 1)
pts = [pts[min(len(pts) - 1, int(round(i * step)))] for i in range(n)]
speed = max(0.1, float(speed or 1.0))
dur = max(0.05, duration_ms(pts) / 1000.0 / speed)
xy = [[x, y] for x, y, _ in pts]
try:
d.swipe_points(xy, dur)
except Exception as e: # 有的设备/agent 不支持多点注入
_log.warning("swipe_points 回放失败(%s),退回直线滑动", e)
d.swipe(xy[0][0], xy[0][1], xy[-1][0], xy[-1][1], dur)
return {"points": len(pts), "duration": round(dur, 3)}
# ================== getevent 解析(纯函数,好单测) ==================
def parse_getevent(lines, max_x, max_y, screen_w, screen_h):
"""把 getevent 行流解析成若干条轨迹 `[[x, y, t_ms], ...]`(屏幕坐标)。
兼容两套协议:`BTN_TOUCH DOWN/UP`(老)与 `ABS_MT_TRACKING_ID`(新,-1=抬起)。
只认 EV_ABS 的 X/Y;Y 到达时才落一个点(X/Y 是先后到达的两个事件)。
"""
strokes, cur = [], None
lx = ly = 0
t0 = None
for ln in lines or []:
m = _LINE_RE.search(ln)
if not m:
continue
ts, kind, code, val = float(m.group(1)), m.group(2), m.group(3), m.group(4)
if kind == "EV_ABS":
if code == "ABS_MT_POSITION_X":
lx = int(val, 16)
elif code == "ABS_MT_POSITION_Y":
ly = int(val, 16)
if cur is not None:
if t0 is None:
t0 = ts
cur.append([lx, ly, (ts - t0) * 1000.0])
elif code == "ABS_MT_TRACKING_ID":
if val.lower().endswith("ffffffff"): # 抬起
if cur:
strokes.append(cur)
cur = None
elif cur is None: # 按下
cur = []
elif kind == "EV_KEY" and code == "BTN_TOUCH":
if val == "DOWN":
if cur is None:
cur = []
else:
if cur:
strokes.append(cur)
cur = None
if cur:
strokes.append(cur)
sx = float(screen_w) / max_x if max_x else 1.0
sy = float(screen_h) / max_y if max_y else 1.0
out = []
for st in strokes:
pts = [[int(round(x * sx)), int(round(y * sy)), t] for x, y, t in st]
pts = clean_points(pts)
if len(pts) >= MIN_POINTS:
out.append(pts)
return out
def find_touch_device(serial, screen_w=0, screen_h=0):
"""找真触屏设备,返回 (device_path, max_x, max_y);找不到返回 (None, 0, 0)。
设备名里带 uinput/virtual/sar/gpio 的一律排除(那些是合成输入);
多个候选时优先名字像触屏的,其次挑坐标范围最接近屏幕的那个。
"""
from core.adb_helper import _adb
out = _adb("-s", serial, "shell", "getevent -pl") or ""
cands, dev, name = [], None, ""
for ln in out.splitlines():
s = ln.strip()
if s.startswith("add device"):
dev, name = s.split(":", 1)[-1].strip(), ""
elif s.startswith("name:"):
name = s.split(":", 1)[1].strip().strip('"')
elif "ABS_MT_POSITION_X" in s or "ABS_MT_POSITION_Y" in s:
m = re.search(r"max (\d+)", s)
if not (dev and m):
continue
key = "x" if "POSITION_X" in s else "y"
for c in cands:
if c[0] == dev:
c[1 if key == "x" else 2] = int(m.group(1))
break
else:
mx = int(m.group(1)) if key == "x" else 0
my = int(m.group(1)) if key == "y" else 0
cands.append([dev, mx, my, name])
real = [c for c in cands
if not any(h in (c[3] or "").lower() for h in _VIRTUAL_HINTS)]
pool = real or cands
if not pool:
return None, 0, 0
def score(c):
hit = any(h in (c[3] or "").lower() for h in _TOUCH_HINTS)
return (0 if hit else 1, abs(c[1] - (screen_w or c[1])) + abs(c[2] - (screen_h or c[2])))
pool.sort(key=score)
best = pool[0]
_log.info(f"[{serial}] 录制用手势设备: {best[0]} ({best[3]}) {best[1]}x{best[2]}"
f",候选 {[c[0] + '/' + (c[3] or '?') for c in cands]}")
return best[0], best[1], best[2]
# ================== 手机端录制 ==================
class PhoneRecorder:
"""在**手机真触屏**上录手指轨迹(getevent 流式读取)。"""
def __init__(self, serial, screen_w=0, screen_h=0):
self.serial = serial
self.screen_w = int(screen_w or 0)
self.screen_h = int(screen_h or 0)
self.device = None
self.max_x = self.max_y = 0
self._proc = None
self._lines = []
self._thread = None
self._stop = threading.Event()
self._started = 0.0
self.error = ""
def start(self):
"""起 getevent 子进程并开始收行。失败抛异常(调用方转成用户可读错误)。"""
self.device, self.max_x, self.max_y = find_touch_device(
self.serial, self.screen_w, self.screen_h)
if not self.device or not self.max_x:
raise RuntimeError("没找到触屏设备节点,无法在手机上录制")
self._proc = subprocess.Popen(
[ADB_PATH, "-s", self.serial, "shell", f"getevent -lt {self.device}"],
stdout=subprocess.PIPE, stderr=subprocess.STDOUT)
self._started = time.time()
self._thread = threading.Thread(target=self._pump, daemon=True)
self._thread.start()
_log.info(f"[{self.serial}] 手机录制已开始({self.device})")
def _pump(self):
try:
for raw in iter(self._proc.stdout.readline, b""):
if self._stop.is_set():
break
self._lines.append(raw.decode("utf-8", "replace").rstrip())
if len(self._lines) > 20000: # 防爆:留最近 2 万行足够解析
del self._lines[:5000]
except Exception as e:
self.error = f"读取事件流失败: {e}"
def stop(self):
"""停止并返回 `[[[x,y,t_ms], ...], ...]`(每条手势一段)。"""
self._stop.set()
try:
if self._proc:
self._proc.terminate()
try:
self._proc.wait(timeout=3)
except Exception:
self._proc.kill()
except Exception:
pass
if self._thread:
self._thread.join(timeout=2)
w = self.screen_w or self.max_x
h = self.screen_h or self.max_y
gestures = parse_getevent(self._lines, self.max_x, self.max_y, w, h)
_log.info(f"[{self.serial}] 手机录制结束:{len(self._lines)} 行事件 → "
f"{len(gestures)} 条手势({sum(len(g) for g in gestures)} 点)")
return gestures