fix: OCR 文字点击落点精度——命中文本块较长时(一行含多段文字)按关键词在文本中的位置比例估算 x,避免点整块中心偏离关键词(tap_text 与任务 OCR 条件共用受益)

This commit is contained in:
2026-09-04 14:33:28 +08:00
parent 44a9dafcd8
commit 4d4f0ae666
+15 -4
View File
@@ -69,11 +69,22 @@ def recognize(image):
def find_on_screen(image, keyword):
"""在截图上查找包含 keyword 的文字。
返回 (found, center_xy, matched_text);center_xy 为文字中心像素坐标(可点击),
未命中返回 (False, None, "")。
返回 (found, center_xy, matched_text);center_xy 为**关键词**中心像素坐标
(可点击),未命中返回 (False, None, "")。
OCR 一行常含多段文字(如 "医值得推荐#苏州济世璞真…展开"),若直接点整块
中心会偏离关键词很远。命中块文本较长时,按关键词在文本中的位置比例估算 x
(块内文字近似等宽),y 取块中心——中文场景估算偏差小,点击能落准。
"""
for r in recognize(image):
if keyword in r["text"]:
text = r["text"]
if keyword in text:
b = r["box"]
return True, ((b[0] + b[2]) // 2, (b[1] + b[3]) // 2), r["text"]
if len(text) <= 12: # 短文本:整块中心即关键词中心
return True, ((b[0] + b[2]) // 2, (b[1] + b[3]) // 2), text
idx = text.find(keyword)
ratio = (idx + len(keyword) / 2) / len(text)
cx = b[0] + int((b[2] - b[0]) * ratio)
cy = (b[1] + b[3]) // 2
return True, (cx, cy), text
return False, None, ""