这份文档是当前后端实现路径的主说明。它把上传链路、fast-first 批改策略、OCR/切题、Mark Scheme、verifier、推荐练习、反馈埋点和 benchmark 结果放在一处,方便后续开发、复测和上线前审查。
A-Level Assistant 不是“上传图片后等一个大模型回答”的系统。当前实现是一条可观测的学习闭环:
上传
-> 可选预处理 / 缓存
-> 流式批改入口
-> OCR + Vision 切题
-> 真题 / Mark Scheme 上下文
-> 快速首轮批改
-> 确定性规则校验
-> SSE 逐题返回
-> 讲解 / 推荐练习
-> 反馈与效果指标
设计取舍是:先返回可信首轮结果,但不把不确定题伪装成确定结论。 慢题、低置信题、识别超时题和空白/答案页-only 样本应显示为 needs_review 或 timeout placeholder。
flowchart TD
U["学生上传图片或 PDF"] --> F["前端上传组件"]
F --> P{"文件类型判断"}
P -->|"图片 / 多图"| C["/prepare-upload<br/>内容哈希缓存<br/>进行中请求去重"]
C -->|"upload_ids"| S["/analyze-homework-stream<br/>流式批改入口"]
P -->|"Large PDF"| L["/large-pdf/prepare<br/>缩略图 + 选页"]
L --> S
S --> Q["快速首题批改<br/>OCR 证据 + 题目识别 + Mark Scheme / 开放批改 + 规则校验"]
Q --> E["SSE 逐题返回<br/>先出首个可信结果"]
E --> T{"还有慢题或高风险题吗"}
T -->|"有"| N["needs_review 占位<br/>不硬编分数"]
T -->|"无"| Z["summary 事件"]
N --> Z
Z --> A["汇总错因与薄弱点<br/>priority_topics / tags / review_count"]
A --> M{"能可靠识别练习主题吗"}
M -->|"不能"| X["recommendation_mode=none<br/>不推荐,提示补充题目来源"]
M -->|"有风险或自定义题"| Y["recommendation_mode=ask_first<br/>先问是否继续练这个点"]
M -->|"真题高置信 / 已确认"| R["recommendation_mode=auto<br/>进入题库查询"]
Y -->|"学生确认"| R
R --> B{"题库有真实候选题吗"}
B -->|"有"| K["输出真实题库练习<br/>基础 / 巩固 / 真题风格"]
B -->|"没有"| X2["不输出假题<br/>说明题库暂无可用候选"]
K --> G["学生开始练习"]
G --> H["提交答案并再次批改"]
H --> I["根据结果调整下一题难度"]
X --> J["反馈埋点与效果看板"]
X2 --> J
Y --> J
K --> J
H --> J
I --> J
当前前端上传路径默认传 fast_batch=true。流程图不再把它画成用户选择,因为真实用户不会主动声明“我要快速首题”;他只会上传图片。质量控制发生在快速路径内部:OCR 只作为证据,规则校验负责兜底,慢题或高风险题用 needs_review 显式暴露,不为了速度硬编分数。
图片和 PDF 的用户意图不同。图片上传通常是“我现在要快点知道这页作业哪里错了”;Large PDF 通常是“我上传了一整套真题,需要先选页”。所以当前前端在进入批改时默认走快速首题返回;PDF 的特殊点不在于慢速批改,而在于先做缩略图和选页,避免把封面、空白页、答案页一起送进批改。
上传链路很容易重复提交同一张图:网络抖动、用户重复点击、前端重试都会发生。内容哈希缓存可以让同图复用识别结果,in-flight dedupe 可以让正在识别的同一张图只跑一次。这两个设计不是为了“炫技”,而是直接减少等待时间和模型成本。
快速首题模式对应 fast_batch=true 路径。它解决的是批改链路里的首屏等待问题:一份上传里可能有多道题,其中有的题很快完成,有的题会因为图片质量、题干缺失、跨页上下文、模型响应慢或校验不确定而拖住整页。如果系统等所有题都完整识别、批改、总结完才显示结果,用户会觉得页面卡住。
它解决三件事:
- 先返回:只要有题目完成识别和批改,就通过 SSE 先推给前端,不等整页所有题都结束。
- 控长尾:首个可用结果返回后,剩余慢题只再等待一个短窗口,避免一两道慢题拖住整次会话。
- 不乱判:慢题、低置信题或识别超时题不会硬编分数,而是返回
needs_review=true的占位结果,提示用户复核或重拍。
什么时候要开:当前产品默认开启。因为用户通常不会写 prompt 说明自己要快还是要慢,他只是上传图片;从体验角度看,先返回第一道可信结果几乎总是更符合预期。
什么时候不要开:不是由普通用户决定,而是由内部调用方或未来自动路由决定。典型场景包括:人工抽检、离线 benchmark、需要完整多模型复核、需要完整讲解先生成完再展示、或某类高风险样本被系统判定不能走快路径。
当前要诚实记录一点:系统已经有 fast_batch=true/false 两条运行路径,但还没有一个成熟的“上传前自动判断是否关闭快速首题”的策略。现阶段的风险控制主要发生在快速路径内部:慢题、低置信题、识别超时题会被标成 needs_review,而不是强行给确定分数。后续如果要真正做自动路由,可以把低清晰度、答案页-only、跨页父题缺失、Past Paper 匹配失败、题量过大等信号前置,自动切到质量优先路径。
批改一整页作业时,某一道题可能因为图片模糊、步骤很长或模型响应慢而拖住全页。如果等所有题都完成,用户会觉得系统卡死。SSE 逐题返回让前端先展示已经可靠完成的题,用户可以先看错因和分数,剩余慢题再继续补齐。
快速模式不是无限等。系统在第一题返回后,只给剩余慢题一个额外短窗口;如果还没完成,就返回 needs_review 占位结果。这样做是为了避免一两道慢题把整次体验拖到 1-2 分钟,同时也避免为了“快”而编造批改结论。
OCR 很适合读取打印题干和部分公式,但对手写步骤、倾斜图片、低清晰度照片并不稳定。如果让 OCR 直接主导结构,可能把学生答案当题干,或把错误步骤改成正确表达。因此 OCR 只有在包含真实任务语言时才进入切题提示;否则只保留为审计证据。
A-Level 批改不是只看最终答案,很多题按 method mark、accuracy mark、结论 mark 给分。能匹配真题时,系统优先使用 Mark Scheme 作为评分依据;匹配不到时才回到开放 AI 批改。这个设计的目的,是让批改尽量接近真实阅卷,而不是只像“看起来合理的解释”。
LLM 很会解释,但会算错、漏符号、把等价表达式误判为不等价。SymPy、统计、概率和化简 verifier 负责校准封闭数学事实。它们不是替代老师或模型,而是作为安全网,专门抓那些“语言看起来对、数学事实错”的情况。
教育产品里最危险的不是慢,而是给学生一个高置信但错误的结论。needs_review 是产品层的诚实机制:当题干缺失、答案保真不足、校验不通过或模型信心不足时,系统把风险显式展示出来,让用户知道这题需要复核或重拍。
summary 不是简单收尾。它把逐题结果整理成后续学习闭环需要的结构化信号,包括总题数、正确/错误/未答数量、得分、review_count、常见错误类型、knowledge_tags_summary 和 priority_topics。推荐练习、效果看板、弱点复盘都依赖这个阶段,如果 summary 只写一句泛泛总评,后面的推荐就会变成拍脑袋。
只告诉学生“你错了”不够。有效学习闭环要告诉学生下一步练什么,所以系统会把错因、topic、subtopic、真题上下文和题库结合起来,决定自动推荐、先询问,还是不推荐。题库信号不足时不硬推,是为了避免把学生带到错误练习方向。
真实逻辑是:
- 没有可靠 topic,或者 paper number 不在 P1-P6:
recommendation_mode=none,不推荐。 - 有 topic 但存在复核风险、自定义题或来源不确定:
ask_first,先问学生要不要继续练这个点。 - Past Paper / Mark Scheme 高置信,或学生已确认:
auto,进入题库查询。 - 题库查到真实候选题:输出真实题库练习,按基础、巩固、真题风格分层。
- 题库没有候选题:不生成假题,返回“当前题库没有可用真实候选题”。
这一步的重点是“有题库才推荐,没有题库就诚实不推荐”。它不是一个生成式练习占位符,而是题库驱动的下一步动作。
推荐题不是终点,学生做完之后还要提交答案。系统会调用练习批改接口,返回得分、短反馈、参考答案和评分标准;然后根据本次练习表现调整下一题难度。这样闭环从“看懂错因”继续到“验证是否真的会了”。
这个产品的质量不能只靠主观感觉。上传成功率、首题返回时间、题号召回、判分一致率、needs_review 命中率、推荐真实题占比、练习开始率、练习提交率、下一题继续率都需要持续追踪。benchmark 和 feedback dashboard 的作用是让每次优化都有证据:到底更快了、推荐是否真来自题库、学生有没有继续练,还是只是把失败藏起来了。
| UI surface | Main code | Runtime behavior |
|---|---|---|
| Upload shell | frontend/src/components/UploadForm.tsx |
Select images/PDF, prepare uploads, submit to streaming analyze |
| Large PDF | frontend/src/api/largePdfClient.ts |
Prepare PDF, show thumbnails, submit selected pages |
| Result page | frontend/src/components/QuestionCard.tsx, PageSummary.tsx |
Render streamed question results, confidence, feedback, explanations |
| Practice loop | frontend/src/components/practice/PracticeRecommendations.tsx |
Ask-first or auto recommendations from questionbank |
| Feedback events | frontend/src/api/client.ts, api/feedback.py |
Track UI funnel and benchmark metadata |
Development server:
python -m uvicorn api.app:app --host 0.0.0.0 --port 8000
cd frontend && npm run devVite proxies upload and practice routes to localhost:8000.
Primary files:
api/routes.pyapi/upload_cache.pypipeline/pipeline.pypipeline/segmenter.pyutils/image_utils.py
Path:
UploadForm
-> /prepare-upload
-> upload_cache stores extracted questions by upload_id/content hash
-> /analyze-homework-stream with fast_batch=true
-> run_pipeline_streaming(..., fast_batch=True)
Important behavior:
- Single image benchmark now matches the real product path by sending
fast_batch=true. api/upload_cache.pyadds content hash cache and in-flight dedupe.- Prepared upload results are reused when healthy; recognition timeout cache entries force full fallback instead of silently dropping pages.
- Fast-first mode returns question events as soon as they are ready.
- After the first question has been emitted, pending slow questions get only a short extra window before returning
needs_reviewtimeout placeholders.
Key knobs:
| Setting | Default | Purpose |
|---|---|---|
PREPARE_UPLOAD_TIMEOUT_SECONDS |
15 |
Upload-time pre-recognition cap per image, far below the old 120s tail |
FAST_BATCH_RECOGNITION_TIMEOUT_SECONDS |
15 |
Click-to-grade recognition cap |
FAST_BATCH_PREPARE_TIMEOUT_SECONDS |
15 |
Fast batch page recognition cap |
FAST_BATCH_IMAGE_MAX_DIMENSION |
1600 |
Long-edge resize for interactive recognition |
FAST_BATCH_QUESTION_TIMEOUT_SECONDS |
15 |
Interactive grading cap per question, far below the old 120s tail |
FAST_BATCH_AFTER_FIRST_QUESTION_TIMEOUT_SECONDS |
8 |
Extra wait after first usable question before pending items are considered timed out |
FAST_BATCH_TIMEOUT_GRACE_SECONDS |
2 |
Short drain window for near-complete fast-batch grading results |
FAST_BATCH_PREPARE_MAX_WORKERS |
10 |
Prepare/upload recognition concurrency |
FAST_BATCH_MAX_WORKERS |
16 |
Fast-first question grading concurrency |
Primary files:
api/routes.pyfrontend/src/api/largePdfClient.tsfrontend/src/components/UploadForm.tsx
Path:
PDF upload
-> /large-pdf/prepare
-> thumbnail + default page selection
-> /large-pdf/{pdf_id}/analyze-stream
-> run_pipeline_streaming on selected rendered pages
The frontend keeps PDF selection explicit, because a large past paper often contains covers, blank pages, mark schemes or answer-only pages.
Primary files:
pipeline/segmenter.pyrouter/models.pyutils/image_utils.py
Current OCR strategy:
- Mathpix is an evidence layer, not the structure authority.
- OCR text enters the segmenter prompt only when the OCR guard sees task language.
- Handwriting-only/formula-only OCR is retained as audit evidence.
- Local tesseract remains a weak fallback and page-header probe.
This prevents handwritten work from overwriting the actual question structure. See Model Routing And OCR Chain.
Primary files:
api/paper_resolver.pyquestionbank/pastpaper_matcher.pyquestionbank/mark_scheme.pypipeline/pipeline.py
Path:
upload intent + file/header/page clues
-> paper resolver
-> questionbank / mark scheme lookup
-> attach mark_scheme_context per question
-> grader uses context when confidence is sufficient
The resolver should never invent certainty. Medium/low confidence flows should ask for confirmation or fall back to open grading.
Primary files:
grader/grader.pypipeline/pipeline.pyverifier/statistics_verifier.pyverifier/math_verifier.pyverifier/probability_verifier.pyverifier/simplification_verifier.py
Current strategy:
- First-pass grading is fast and streamed.
- Risky cases use review/multi-agent paths when enabled.
- Deterministic verifiers can correct or flag LLM grading for closed-form math.
needs_review=trueis preferred over a confident wrong answer.
Fast-first mode deliberately disables inline solution generation and heavy summary LLM calls; deeper explanation remains available through follow-up actions.
Primary files:
api/practice_orchestrator.pyquestionbank/database.pyfrontend/src/components/practice/PracticeRecommendations.tsx
Inputs:
priority_topicsknowledge_tags_summary- wrong/unanswered questions
- upload intent and paper context
Outputs:
auto: enough topic/paper confidence, return real questionbank items.ask_first: topic is plausible but needs student confirmation.none: insufficient signal.
The evaluator accepts narrow topic aliases such as sigma_notation within the statistics group to avoid false negative recommendation scores.
Primary files:
api/feedback.pyapi/effectiveness.pyscripts/evaluate_upload_corpus.pyscripts/build_jpeg_benchmark_corpus.py
Important events/metrics:
| Metric | Meaning |
|---|---|
upload_success_rate |
Request completed successfully |
parse_success_rate |
At least one readable question extracted |
readable_question_rate |
Question-level readable/non-timeout rate |
first_question_p95_ms |
User-facing first result latency |
image_end_to_end_p95_ms |
Full image session latency |
sse_first_event_p95_ms |
Transport/proxy responsiveness |
segmentation_done_p95_ms |
Recognition/cutting latency |
first_grading_after_segmentation_p95_ms |
First grading latency after segmentation |
summary_after_first_question_p95_ms |
How long slow remaining questions delay the page |
recognition_timeout_count |
Recognition hard failures |
fast_batch_timeout_count |
Question grading hard fallbacks |
marked_correctness_match_rate |
Labeled quality benchmark correctness |
recommendation_relevance_rate |
Labeled recommendation topic relevance |
Latest focused reports:
reports/effectiveness/20260626_fast_first_single_image_report.mdreports/effectiveness/20260626_jpeg30_phase_benchmark_report.mdreports/effectiveness/20260626_quality_speed_iteration_report.md
Stable wins:
| Area | Latest evidence |
|---|---|
| 10-image prepared batch | Overall pass 100, 20/20 correct, first question 4.244s |
| Product-path single JPEG | Overall pass 100 on seed corpus, first question 24.618s |
| Fixed 30-image corpus | Upload success 100%, parse success 100%, readable question rate 100% |
| Backend focused regression | 72 passed, 1 skipped |
| Frontend build | Passes; Vite chunk-size warning remains |
Known failures from JPEG30:
| Metric | Result | Target |
|---|---|---|
image_end_to_end_p95_ms |
131.071s | 60s |
first_question_p95_ms |
88.901s | 30s |
segmentation_done_p95_ms |
77.197s | 20s |
first_grading_after_segmentation_p95_ms |
34.809s | 15s |
summary_after_first_question_p95_ms |
113.997s before the short-window fix | 15s |
fast_batch_timeout_count |
2 | 0 |
Interpretation:
- SSE/proxy is fast; the bottleneck is not transport.
- Tilted/shadow and cross-page photos are recognition long-tail cases.
- Blank/answer-only images should be rule-detected earlier instead of sent through full grading.
- Before the latest short-window fix, later slow questions could delay the page even when the first question was already available.
Priority order:
- Keep fast-first short-window behavior: after the first question, do not let slow pending questions block the whole page for 120s.
- Add blank/answer-only lightweight detection before grading.
- Add image quality scoring for tilt, shadow, crop risk and show retake guidance.
- Add repeat benchmark for the 30-image corpus after each performance change.
- Label the 30-image corpus with expected question count/order/score so it becomes both a speed and quality benchmark.
- Split frontend bundles, especially PDF worker and non-first-screen demo/practice code.
Focused backend:
PYTHONPATH=. pytest -q \
test/test_pipeline_streaming.py \
test/test_fast_upload_flow.py \
test/test_effectiveness.py \
test/test_large_pdf_mode.pyBroader backend set used in recent iterations:
PYTHONPATH=. pytest -q \
test/test_statistics_verifier.py \
test/test_fast_upload_flow.py \
test/test_pipeline_streaming.py \
test/test_effectiveness.py \
test/test_large_pdf_mode.py \
test/test_practice_orchestrator.py \
test/test_rescue_bridging.py \
test/test_feedback_metrics_dashboard.pyFrontend:
cd frontend
npm run buildBenchmark corpus:
python scripts/build_jpeg_benchmark_corpus.py \
--output-dir test/fixtures/jpeg_benchmark_corpus \
--count 30
python scripts/evaluate_upload_corpus.py \
--input-dir test/fixtures/jpeg_benchmark_corpus \
--api-base http://127.0.0.1:8000 \
--repeat 1 \
--max-concurrency 1 \
--track-events \
--output reports/effectiveness/jpeg30_fast_first_phase_metrics_YYYYMMDD.json- Do not let OCR unconditionally rewrite segmenter output.
- Do not hide uncertainty by dropping timeout/unreadable placeholder questions.
- Do not make fast-first generate inline solutions before first result.
- Do not widen topic aliases without a matching benchmark sample or human review note.
- Do not optimize speed by lowering
needs_reviewvisibility.