win b27686a409 docs(01-shared-core): create phase 1 plans for crawler_core shared package
Plan 01-01 (Wave 1): Package scaffold with HTTPClient + tenacity retry (min=10s)
+ stdlib logging + BaseFetcher/BaseSearcher base classes + pyproject.toml.
Covers ARCH-01, ARCH-02, QUAL-04, QUAL-05.

Plan 01-02 (Wave 2): Sign algorithm migration (Boss/Job51/Zhilian) to
crawler_core/ + comprehensive unit tests — no HTTP, no mocks, pure functions.
Covers QUAL-01. 24+ test cases across 3 test files.

ROADMAP updated: Phase 1 now shows 2 concrete plans instead of TBD.
2026-03-21 17:45:14 +08:00

517 lines
20 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
phase: 01-shared-core
plan: 01
type: execute
wave: 1
depends_on: []
files_modified:
- crawler_core/__init__.py
- crawler_core/http_client.py
- crawler_core/base.py
- crawler_core/boss/__init__.py
- crawler_core/qcwy/__init__.py
- crawler_core/zhilian/__init__.py
- crawler_core/pyproject.toml
- Pipfile
autonomous: true
requirements:
- ARCH-01
- ARCH-02
- QUAL-04
- QUAL-05
must_haves:
truths:
- "`pip install -e ./crawler_core` succeeds without errors"
- "`from crawler_core import BaseFetcher, BaseSearcher, ApiResult, HTTPClient` imports cleanly"
- "HTTPClient retries failed requests up to 3 times with exponential backoff (minimum 10s wait)"
- "All HTTP errors are logged to stderr via stdlib logging with level, url, and error message"
- "Old spiderJobs/ and jobs_spider/ code is NOT modified — feature flag isolation holds"
artifacts:
- path: "crawler_core/pyproject.toml"
provides: "Package metadata for editable install"
contains: "name = \"crawler_core\""
- path: "crawler_core/__init__.py"
provides: "Public API surface"
exports: ["BaseFetcher", "BaseSearcher", "ApiResult", "HTTPClient"]
- path: "crawler_core/http_client.py"
provides: "TLS-fingerprinted HTTP client with retry and logging"
exports: ["HTTPClient"]
- path: "crawler_core/base.py"
provides: "Template-method base classes"
exports: ["ApiResult", "BaseFetcher", "BaseSearcher", "parse_response"]
key_links:
- from: "crawler_core/__init__.py"
to: "crawler_core/http_client.py"
via: "from crawler_core.http_client import HTTPClient"
- from: "crawler_core/__init__.py"
to: "crawler_core/base.py"
via: "from crawler_core.base import BaseFetcher, BaseSearcher, ApiResult"
- from: "crawler_core/base.py"
to: "crawler_core/http_client.py"
via: "from crawler_core.http_client import HTTPClient"
---
<objective>
Create the crawler_core/ installable shared package with its core infrastructure: HTTP client with TLS fingerprint, retry logic, stdlib logging, and the BaseFetcher/BaseSearcher template-method base classes.
Purpose: This is the foundation everything else depends on. Once installed with `pip install -e ./crawler_core`, Phase 2/3 platform rewrites can import from it instead of copying code.
Output: A working Python package at crawler_core/ that installs cleanly and exposes BaseFetcher, BaseSearcher, ApiResult, and HTTPClient.
</objective>
<execution_context>
@~/.claude/get-shit-done/workflows/execute-plan.md
@~/.claude/get-shit-done/templates/summary.md
</execution_context>
<context>
@.planning/PROJECT.md
@.planning/ROADMAP.md
@.planning/phases/01-shared-core/1-CONTEXT.md
<interfaces>
<!-- Key existing code the executor must understand before creating crawler_core/ equivalents. -->
<!-- DO NOT copy these verbatim — update the internal import paths. -->
From spiderJobs/core/http_client.py:
```python
class HTTPClient:
def __init__(self, base_url, default_headers=None, proxy=None,
tunnel_proxy=None, proxy_pool=None, timeout=10): ...
def _new_session(self) -> requests.Session: ...
def _get_proxies(self) -> Optional[dict]: ...
def _merge_headers(self, extra=None) -> dict: ...
def post(self, path, body, headers=None) -> tuple[int, Any]: ...
def get(self, path, params=None, headers=None) -> tuple[int, Any]: ...
```
From spiderJobs/core/base.py:
```python
@dataclass
class ApiResult:
success: bool
status_code: int
data: Any = None
list: list[dict] = field(default_factory=list)
count: int = 0
is_end_page: bool = True
error: Optional[str] = None
def parse_response(http_code: int, raw: Any) -> ApiResult: ...
class BaseFetcher:
ENDPOINT: str = ""
def __init__(self, http_client: HTTPClient): ...
def _build_params(self) -> dict: raise NotImplementedError
def _parse(self, http_code, raw) -> ApiResult: ...
def fetch(self) -> ApiResult: ...
class BaseSearcher:
ENDPOINT: str = ""
def __init__(self, page_size=15, http_client=None): ...
def _build_params(self, page_index) -> dict: raise NotImplementedError
def _request(self, params) -> tuple[int, Any]: ...
def _parse(self, http_code, raw) -> ApiResult: ...
def search(self, page_index=1) -> ApiResult: ...
def load_all(self, max_pages=10, on_page=None) -> list[dict]: ...
```
</interfaces>
</context>
<tasks>
<task type="auto">
<name>Task 1: Create crawler_core package scaffold and pyproject.toml</name>
<read_first>
- /Users/win/2025/AICoding/JobData/pyproject.toml (understand existing project config format)
- /Users/win/2025/AICoding/JobData/Pipfile (understand dependency structure to add entries)
- /Users/win/2025/AICoding/JobData/.planning/phases/01-shared-core/1-CONTEXT.md (decisions D-01 through D-04)
</read_first>
<files>
crawler_core/pyproject.toml
crawler_core/boss/__init__.py
crawler_core/qcwy/__init__.py
crawler_core/zhilian/__init__.py
Pipfile
</files>
<action>
Create the crawler_core/ directory structure and configure it as an installable Python package.
**Step 1: Create crawler_core/pyproject.toml**
```toml
[build-system]
requires = ["setuptools>=68"]
build-backend = "setuptools.backends.legacy:build"
[project]
name = "crawler_core"
version = "0.1.0"
description = "Shared crawler core — sign algorithms, HTTP client, base classes"
requires-python = ">=3.11"
dependencies = [
"requests_go==1.0.9",
"tenacity>=8.0",
]
[tool.setuptools.packages.find]
where = [".."]
include = ["crawler_core*"]
```
NOTE: `where = [".."]` means setuptools finds the `crawler_core` package by looking one level up from the pyproject.toml, which is at the repo root. This makes `pip install -e ./crawler_core` resolve correctly.
**Step 2: Create platform namespace __init__.py files (empty)**
Create these three files with a single docstring only — NO imports, they are just namespace markers:
- `crawler_core/boss/__init__.py`: `"""Boss直聘 platform module."""`
- `crawler_core/qcwy/__init__.py`: `"""前程无忧 (51Job) platform module."""`
- `crawler_core/zhilian/__init__.py`: `"""智联招聘 platform module."""`
**Step 3: Add dependencies to Pipfile**
In the `[packages]` section (before `[dev-packages]`), add these two lines (after `playwright = "==1.57.0"`):
```
requests_go = "==1.0.9"
tenacity = ">=8.0"
```
In the `[dev-packages]` section, add:
```
pytest = ">=8.0"
pytest-cov = ">=4.0"
pytest-anyio = "*"
```
**What NOT to do:**
- Do NOT create a crawler_core/__init__.py in this task (Task 2 creates it)
- Do NOT create crawler_core/http_client.py or crawler_core/base.py (Task 2 and 3)
- Do NOT run `pip install` — just write the files
</action>
<verify>
<automated>python -c "import tomllib; d=tomllib.load(open('/Users/win/2025/AICoding/JobData/crawler_core/pyproject.toml','rb')); assert d['project']['name']=='crawler_core'; print('pyproject.toml OK')" && grep -q "requests_go" /Users/win/2025/AICoding/JobData/Pipfile && grep -q "tenacity" /Users/win/2025/AICoding/JobData/Pipfile && grep -q "pytest" /Users/win/2025/AICoding/JobData/Pipfile && echo "Pipfile OK"</automated>
</verify>
<acceptance_criteria>
- `crawler_core/pyproject.toml` exists and contains `name = "crawler_core"`, `requires-python = ">=3.11"`, `requests_go==1.0.9`, `tenacity>=8.0`
- `crawler_core/boss/__init__.py`, `crawler_core/qcwy/__init__.py`, `crawler_core/zhilian/__init__.py` all exist (can be empty docstrings)
- `Pipfile` [packages] section contains `requests_go = "==1.0.9"` and `tenacity = ">=8.0"`
- `Pipfile` [dev-packages] section contains `pytest`, `pytest-cov`, `pytest-anyio`
- `grep -c "requests_go" /Users/win/2025/AICoding/JobData/Pipfile` outputs `1` (no duplicates)
</acceptance_criteria>
<done>Package directory structure created, pyproject.toml valid, dependencies declared in Pipfile.</done>
</task>
<task type="auto">
<name>Task 2: Create crawler_core/http_client.py with tenacity retry and logging</name>
<read_first>
- /Users/win/2025/AICoding/JobData/spiderJobs/core/http_client.py (source to port — read every line)
- /Users/win/2025/AICoding/JobData/.planning/research/STACK.md (tenacity config section, TLS fingerprint section)
- /Users/win/2025/AICoding/JobData/.planning/phases/01-shared-core/1-CONTEXT.md (D-03: no loguru, stdlib only; D-09: one HTTPClient class)
</read_first>
<files>
crawler_core/http_client.py
</files>
<action>
Port `spiderJobs/core/http_client.py` to `crawler_core/http_client.py` with two additions: tenacity retry and stdlib logging.
**The file must be exactly `crawler_core/http_client.py`** — no subdirectory.
**Imports to use (CRITICAL — per D-03, only requests_go + stdlib + tenacity):**
```python
from __future__ import annotations
import logging
import random
from typing import Any, Optional
import requests_go as requests
from requests_go.tls_config import TLS_CHROME_LATEST
from tenacity import (
retry,
retry_if_exception_type,
stop_after_attempt,
wait_random_exponential,
)
```
**Logging setup (module-level, before the class):**
```python
logger = logging.getLogger("crawler_core.http_client")
```
This uses stdlib logging — NOT loguru (per D-03, loguru is excluded from crawler_core). Callers (app/services/crawler/) can configure loguru to bridge stdlib if desired.
**Class structure:** Copy the full `HTTPClient` class from `spiderJobs/core/http_client.py` EXACTLY, then make these targeted changes:
1. **Keep all existing methods unchanged:** `__init__`, `_new_session`, `_get_proxies`, `_merge_headers`
2. **Wrap `post()` with tenacity retry decorator:**
```python
@retry(
stop=stop_after_attempt(3),
wait=wait_random_exponential(multiplier=1, min=10, max=30),
retry=retry_if_exception_type((ConnectionError, TimeoutError, OSError)),
reraise=True,
before_sleep=lambda retry_state: logger.warning(
"HTTP retry attempt=%d url=%s error=%s",
retry_state.attempt_number,
retry_state.args[1] if retry_state.args else "unknown",
retry_state.outcome.exception(),
),
)
def post(self, path: str, body: dict, headers: Optional[dict] = None) -> tuple[int, Any]:
"""发送 POST 请求"""
# ... existing body unchanged ...
logger.debug("POST %s%s", self.base_url, path)
# existing try/finally logic unchanged
```
3. **Wrap `get()` with the same tenacity retry decorator** (identical decorator, same pattern):
```python
@retry(
stop=stop_after_attempt(3),
wait=wait_random_exponential(multiplier=1, min=10, max=30),
retry=retry_if_exception_type((ConnectionError, TimeoutError, OSError)),
reraise=True,
before_sleep=lambda retry_state: logger.warning(
"HTTP retry attempt=%d error=%s",
retry_state.attempt_number,
retry_state.outcome.exception(),
),
)
def get(self, path: str, params: Optional[dict] = None, headers: Optional[dict] = None) -> tuple[int, Any]:
"""发送 GET 请求"""
logger.debug("GET %s%s", self.base_url, path)
# ... existing body unchanged ...
```
4. **Add module docstring at the top:**
```python
"""
crawler_core.http_client — 通用 HTTP 客户端
基于 requests-go自带 Chrome TLS 指纹伪装TLS_CHROME_LATEST + random_ja3=True
支持代理 IP / 隧道代理 / 代理池轮换。
内置 tenacity 重试3次指数退避最小10秒间隔
使用 stdlib logging — 上层可通过 logging.getLogger('crawler_core') 配置。
不依赖 loguru / FastAPI / Tortoise-ORM 等应用框架。
"""
```
**Minimum 10 second wait is MANDATORY**`min=10` in `wait_random_exponential` preserves the anti-detection delay requirement from STACK.md.
**Do NOT:**
- Change the proxy logic (keep tunnel_proxy / proxy_pool / fixed proxy logic identical)
- Import loguru
- Import anything from `spiderJobs.*` or `app.*`
</action>
<verify>
<automated>cd /Users/win/2025/AICoding/JobData && python -c "
import sys
sys.path.insert(0, '.')
from crawler_core.http_client import HTTPClient
import inspect, logging
src = inspect.getsource(HTTPClient.post)
assert 'retry' in src or '@retry' in dir(HTTPClient.post), 'tenacity decorator missing on post'
assert 'logger' in src or 'logging' in src, 'logging missing in post'
print('HTTPClient OK')
"</automated>
</verify>
<acceptance_criteria>
- `crawler_core/http_client.py` exists and is importable: `from crawler_core.http_client import HTTPClient` succeeds (after adding crawler_core to sys.path)
- File contains `from tenacity import retry` in imports
- File contains `logger = logging.getLogger("crawler_core.http_client")`
- File contains `wait_random_exponential(multiplier=1, min=10, max=30)` — exact values
- File contains `stop_after_attempt(3)`
- File does NOT contain `import loguru` or `from loguru` anywhere
- File does NOT contain `from spiderJobs` or `from app` anywhere
- File is under 200 lines (source is 155 lines + ~30 lines of additions)
- `grep -c "from tenacity" /Users/win/2025/AICoding/JobData/crawler_core/http_client.py` outputs `1`
- `grep "min=10" /Users/win/2025/AICoding/JobData/crawler_core/http_client.py` has output
</acceptance_criteria>
<done>HTTPClient ported with retry (3 attempts, min=10s wait) and stdlib logging. No loguru. No spiderJobs imports.</done>
</task>
<task type="auto">
<name>Task 3: Create crawler_core/base.py and crawler_core/__init__.py</name>
<read_first>
- /Users/win/2025/AICoding/JobData/spiderJobs/core/base.py (source to port — read every line)
- /Users/win/2025/AICoding/JobData/.planning/research/ARCHITECTURE.md (abstract base class hierarchy section)
- /Users/win/2025/AICoding/JobData/.planning/phases/01-shared-core/1-CONTEXT.md (D-05, D-06, D-07: base class interface decisions)
</read_first>
<files>
crawler_core/base.py
crawler_core/__init__.py
</files>
<action>
Port `spiderJobs/core/base.py` to `crawler_core/base.py` and create the public `__init__.py`.
**crawler_core/base.py:**
Port the full file from `spiderJobs/core/base.py` with ONE import change:
Change:
```python
from spiderJobs.core.http_client import HTTPClient
```
To:
```python
from crawler_core.http_client import HTTPClient
```
Everything else stays identical to `spiderJobs/core/base.py`:
- `ApiResult` dataclass with fields: `success`, `status_code`, `data`, `list`, `count`, `is_end_page`, `error`
- `parse_response(http_code, raw)` function
- `BaseFetcher` class with `ENDPOINT`, `__init__`, `_build_params`, `_parse`, `fetch`
- `BaseSearcher` class with `ENDPOINT`, `__init__`, `_build_params`, `_request`, `_parse`, `search`, `load_all`
Add module docstring at the top:
```python
"""
crawler_core.base — 通用基类与数据结构
提供所有招聘平台共用的: ApiResult, BaseFetcher, BaseSearcher, parse_response
不依赖任何平台特定代码。
"""
```
Replace the existing inline print in `load_all`:
```python
# Change this:
print(f"第 {page_index} 页失败: {result.error}")
# To this (use stdlib logging, not print):
import logging as _logging
_log = _logging.getLogger("crawler_core.base")
_log.warning("第 %d 页失败: %s", page_index, result.error)
```
Actually, define the logger at module level (not inside the method):
```python
import logging
_logger = logging.getLogger("crawler_core.base")
```
Then in `load_all`, replace `print(...)` with `_logger.warning(...)`.
**crawler_core/__init__.py:**
```python
"""
crawler_core — 招聘爬虫共享核心包
安装方式: pip install -e ./crawler_core
使用方式: from crawler_core import BaseFetcher, BaseSearcher, ApiResult, HTTPClient
"""
from crawler_core.base import ApiResult, BaseFetcher, BaseSearcher, parse_response
from crawler_core.http_client import HTTPClient
__all__ = [
"ApiResult",
"BaseFetcher",
"BaseSearcher",
"HTTPClient",
"parse_response",
]
__version__ = "0.1.0"
```
**Do NOT:**
- Change the logic of `BaseFetcher.fetch()`, `BaseSearcher.search()`, or `BaseSearcher.load_all()` beyond the logger swap
- Import from `spiderJobs.*` or `app.*`
- Import loguru
- Add any platform-specific code to base.py or __init__.py
</action>
<verify>
<automated>cd /Users/win/2025/AICoding/JobData && python -c "
import sys
sys.path.insert(0, '.')
from crawler_core import BaseFetcher, BaseSearcher, ApiResult, HTTPClient, parse_response
import dataclasses
fields = {f.name for f in dataclasses.fields(ApiResult)}
assert fields == {'success','status_code','data','list','count','is_end_page','error'}, f'ApiResult fields wrong: {fields}'
assert hasattr(BaseFetcher, 'fetch'), 'BaseFetcher.fetch missing'
assert hasattr(BaseSearcher, 'load_all'), 'BaseSearcher.load_all missing'
print('All imports OK, ApiResult fields OK')
"</automated>
</verify>
<acceptance_criteria>
- `from crawler_core import BaseFetcher, BaseSearcher, ApiResult, HTTPClient` succeeds (with repo root on sys.path)
- `crawler_core/base.py` does NOT contain `from spiderJobs` anywhere: `grep "from spiderJobs" crawler_core/base.py` returns empty
- `crawler_core/base.py` does NOT contain `print(` anywhere: `grep "print(" crawler_core/base.py` returns empty
- `crawler_core/__init__.py` contains `__all__` with all 5 exports
- `crawler_core/__init__.py` contains `__version__ = "0.1.0"`
- `ApiResult` dataclass has exactly 7 fields: success, status_code, data, list, count, is_end_page, error
- `BaseFetcher._build_params` raises `NotImplementedError`
- `BaseSearcher._build_params` raises `NotImplementedError`
</acceptance_criteria>
<done>base.py ported (no spiderJobs imports, no print statements), __init__.py exposes clean public API.</done>
</task>
</tasks>
<verification>
Run the full import chain to verify the package works end-to-end before moving to Plan 02:
```bash
cd /Users/win/2025/AICoding/JobData
python -c "
import sys
sys.path.insert(0, '.')
from crawler_core import BaseFetcher, BaseSearcher, ApiResult, HTTPClient, parse_response
# Verify ApiResult structure
r = ApiResult(success=True, status_code=200)
assert r.success and r.list == [] and r.error is None
# Verify BaseFetcher requires _build_params
class TestFetcher(BaseFetcher):
ENDPOINT = '/test'
def _build_params(self):
return {'q': 'test'}
# Verify parse_response with dict input
result = parse_response(200, {'statusCode': 200, 'data': {'list': [{'id': 1}], 'count': 1, 'isEndPage': False}})
assert result.success
assert result.list == [{'id': 1}]
assert not result.is_end_page
print('All verification checks passed')
"
```
Also confirm no cross-contamination:
```bash
grep -r "from spiderJobs" /Users/win/2025/AICoding/JobData/crawler_core/ && echo "FAIL: found spiderJobs import" || echo "OK: no spiderJobs imports"
grep -r "from app" /Users/win/2025/AICoding/JobData/crawler_core/ && echo "FAIL: found app import" || echo "OK: no app imports"
grep -r "loguru" /Users/win/2025/AICoding/JobData/crawler_core/ && echo "FAIL: found loguru" || echo "OK: no loguru"
```
</verification>
<success_criteria>
1. `python -c "from crawler_core import BaseFetcher, BaseSearcher, ApiResult, HTTPClient"` exits 0 (with repo root on sys.path)
2. `crawler_core/pyproject.toml` passes `python -c "import tomllib; tomllib.load(open('crawler_core/pyproject.toml','rb'))"`
3. `grep "requests_go" Pipfile` has output — dependency declared
4. `grep "tenacity" Pipfile` has output — dependency declared
5. `grep "pytest" Pipfile` has output — dev dependency declared
6. `grep -r "from spiderJobs" crawler_core/` has NO output
7. `grep -r "loguru" crawler_core/` has NO output
8. `grep "min=10" crawler_core/http_client.py` has output — anti-detection delay preserved
9. `spiderJobs/` and `jobs_spider/` directories are UNCHANGED (no files modified)
</success_criteria>
<output>
After completion, create `.planning/phases/01-shared-core/01-01-SUMMARY.md` with:
- What was created (file list with line counts)
- Key decisions made (pyproject.toml structure, tenacity config values, logging approach)
- Interface contracts (the public exports from crawler_core/__init__.py)
- Any deviations from this plan and why
</output>