Tiancheng Lu
← Back home

Private product · sanitized case study · runs daily

Job-opportunity screening Agent with human-in-the-loop

The machine screens and presents evidence; the human decides — a local-first job-opportunity evaluation system where real actions never fire automatically.

Role
Solo developer
Period
2026.06 – 2026.07
Status
Resident service · runs daily
Shape
Local-first · zero cloud dependency

The problem

Job hunting has an awkward middle ground: purely manual — scanning dozens to hundreds of postings a day and judging each by eye is repetitive, subjective, and leaks opportunities; fully automated mass-applying — mistargeted sends, inflated claims, repeated outreach — damages your own professional reputation and crosses platform rule lines.

The right shape is the middle: the machine collects, standardizes, pre-screens, presents evidence and explains gaps; the human makes the final call. Any outward-facing real action (contact, apply) is off by default, requires explicit authorization, and leaves a full audit trail.

System boundary


Source adapter → normalization + content-hash dedup → candidate profile → rule + local-LLM hybrid analysis → human review workbench → (ten-condition safety gate) → first-contact queue
ComponentResponsibility
Source adapter Scheduled scanning + detail enrichment; empty JDs never analyzed; risk-control signals trip a breaker that stops collection
Normalization & dedup job_uid = source + sha256(origin|URL|title|company)[:16], identical algorithm on both collection and service ends
Candidate profile Structured skills / experience fact base — analysis must reference the real profile, never judge from the JD alone
Hybrid analyzer Rule layer (hard-requirement extraction, gap classification, thresholds) + local LLM (evidence, risks, messaging); rule fallback when the model fails
Human review workbench Three panes: pending queue / original JD / evidence & decision; delete = permanent exclusion, skip = soft state
Safety gate A ten-condition serial check before any real action
Ops ledger system_events unified event stream: batches, gate verdicts and skip reasons are all traceable
  • Local-first: local LLM (qwen3-14b-class open models) + SQLite + a resident local service — zero cloud dependency in the core path.

Key decisions

  1. 1

    Analysis-only by default; real actions need multiple explicit grants

    A ten-condition serial gate: master switch → decision whitelist → score above threshold (strictest of policy / requirements / preference) → no unsupported claims → salary met → commute acceptable → zero hard-requirement gaps → no blocking-tag gaps (one-vote veto) → under daily cap → dual-end idempotency → not already in conversation. Any failure skips with the reason written to audit.

  2. 2

    Honesty guardrails are configuration, not promises

    forbidden_claims enumerates "what must never be said + the corrected phrasing" (e.g. never self-describe as an algorithm engineer, never invent project types never done); if analysis output contains unsupported claims it is blocked and returned for human correction.

  3. 3

    Privacy never leaves the machine

    Phone, email, full resume and exact address are listed in never_send_to_external_api; defaulting to a local model makes the privacy constraint hold at the architecture level, not through self-discipline.

  4. 4

    Hard-requirement extraction is a deterministic rule layer

    Extract requirement sections from JD text (recognizing heading variants, truncating "nice-to-have / preferred"), classify by capability group into gap / match / unknown, refuse weak evidence, cap evidence at 2 items. The rule layer is 100% deterministically testable and never drifts with model versions.

  5. 5

    Permanent exclusion vs soft skip, semantically separated

    Manual delete = permanent exclusion (exclusion table, never re-ingested); skip = soft state (re-analysis allowed). Avoids both the "deleted but it keeps coming back" loop and losing something forever by accident.

  6. 6

    One unified ops event ledger

    Batch summaries, gate verdicts, skip reasons and notification attempts all land in system_events, queryable from the workbench — for a one-person operation, observability is the lifeline.

  7. 7

    Product boundary held by three questions

    Every new feature must answer: does it improve the core scan → judge → confirm loop? Does it make the system more likely to lie / misjudge / mis-send? Is it the main loop or a side quest? Fail the first question and it does not get built. Explicitly out of scope: mass applying, automated chat follow-up, generic job-search SaaS.

View sanitized architecture (Mermaid source) ›

The diagram uses generic role labels only — no platform names, API details or real data.

%% 职位机会筛选与人工决策 Agent · 脱敏架构图(公开版)
%% 不含平台名、接口细节、登录态机制、真实数据
flowchart TB
  User([本人 · 决策者])

  subgraph Local["本地环境(本地优先 · 零云依赖)"]
    subgraph Adapter["来源适配器层"]
      Scan["定时扫描器<br/>列表 + 详情补全<br/>空 JD 不分析 · 风控信号熔断"]
    end

    subgraph Core["核心服务 · FastAPI + SQLite"]
      Norm["标准化 + 去重<br/>job_uid 内容哈希<br/>双端同算法"]
      Profile["候选人画像库<br/>结构化事实 · 一等公民"]
      Rules["规则层(确定性)<br/>硬性要求提取 · 三态判定<br/>弱证据不判匹配"]
      LLM["本地 LLM 分析<br/>证据 / 风险 / 缺口 / 话术<br/>失败自动规则兜底"]
      Guard["诚实与隐私护栏<br/>forbidden_claims 纠正口径<br/>隐私字段不出机"]
      Gate["十条件安全闸<br/>阈值 · 诚实 · 硬性缺口 · 阻断标签<br/>日上限 · 双端幂等 · 已沟通"]
      Events["system_events 统一账本<br/>批次 / 闸判定 / 跳过原因可追溯"]
    end

    subgraph UI["人工审阅工作台"]
      Queue["待审队列 / 原始 JD / 证据与决策<br/>删除=永久排除 · 跳过=软状态"]
      Logs["运行日志 / 排除列表 / 配置"]
    end
  end

  Source["职位来源<br/>(当前实例:某主流招聘平台)"]
  Model["本机 LLM 运行时<br/>开源模型 · 隐私不出机"]
  Action["首次接触队列<br/>默认 analysis-only<br/>真实动作需显式授权"]

  Source --> Adapter
  Scan --> Norm
  Norm --> Rules
  Profile --> Rules
  Profile --> LLM
  Rules --> LLM
  LLM --> Guard
  Guard --> Queue
  User --> Queue
  Queue -->|"人工决策:入库 / 跳过 / 准备联系"| Gate
  Gate -->|全部通过| Action
  Gate -->|任一不满足:跳过并记录原因| Events
  Scan --> Events
  Model --> LLM
  Events --> Logs

Verification

  • 28 safety-gate unit tests, all passing: each of the ten gate conditions tested to block individually, plus an all-pass enqueue test and a master-switch short-circuit test; hard-requirement extraction tests for classification / truncation / weak evidence; guardrail config completeness tests (missing config fails the test — a fallen defense must be loudly visible);
  • 20-case synthetic evaluation set, 100% passing: fully synthetic job samples covering the three-state classification, blocking tags, section boundaries, weak evidence and fallback paths; tests only the deterministic rule layer, no LLM involved, reproducible anywhere;
  • Runs daily for real: resident service + scheduled scanning + workbench review is a system in daily use, not a demo;
  • doctor / e2e self-check scripts cover service health and key paths.

Public artifacts

  • The safety-gate test suite (28 cases) and synthetic evaluation set (20 cases) are maintained in-repo; run instructions in the project verification.md

Redaction boundary

The following are intentionally excluded from this case study:

  • Platform name, private APIs and request parameter details;
  • Login-state persistence, cookie handling, anti-risk-control or bypass methods;
  • Real companies, real postings, real contacts;
  • Real resumes and candidate privacy data;
  • Any bulk-contact scripting; demos use analysis-only mode and fully synthetic data.