- 网页爬虫
- 人工智能
- AI 应用
【免费下载链接】Scrapegraph-ai
Python scraper based on AI
本篇技术指南以仓库根目录下的 TESTING_INFRASTRUCTURE.md 为主干,结合
pytest.ini、tests/conftest.py、tests/fixtures/ 与 .github/workflows/test-suite.yml 等真实源码,系统拆解 ScrapeGraphAI 的增强测试体系。你将掌握其单元测试、集成测试、性能基准测试的组织方式,学会使用 Mock HTTP 服务器与共享 fixtures 编写可复现的测试,并理解 GitHub Actions 中从单元测试到覆盖率报告的全自动流水线,从而能够直接在本仓库上运行、扩展与定制测试。
一、测试基础设施总览
ScrapeGraphAI 是一套基于 AI 的 Python 网页抓取框架(核心图实现位于 scrapegraphai/graphs/)。为了让这套以 LLM 为引擎的抓取逻辑在多次迭代中保持可靠,仓库引入了一套全栈式增强测试基础设施,覆盖四个层次:
- 单元测试:借助 Mock 对象与本地 fixtures 快速验证节点、图与工具函数的逻辑,不依赖任何外部网络与付费 API;
- 集成测试:连接真实 LLM 提供商(OpenAI、Ollama、Anthropic、Groq 等)与测试网站,验证端到端抓取能力;
- 性能基准测试:跟踪执行时间、Token 消耗、API 调用次数,实现性能回归检测;
- 自动化 CI/CD:在 Ubuntu、macOS、Windows 三平台与 Python 3.10/3.11/3.12 三版本上自动执行上述全部测试。
这一体系对应的文件布局为:
pytest.ini # pytest 与覆盖率配置(位于仓库根目录) tests/ ├── conftest.py # 共享 fixtures 与自定义 pytest 插件逻辑 ├── fixtures/ │ ├── mock_server/server.py # 本地 Mock HTTP 服务器 │ ├── benchmarking.py # 性能基准框架(数据类 + Tracker + 报告) │ ├── helpers.py # 断言辅助、Mock 响应构建、数据生成器 │ └── __init__.py ├── integration/ # 三类集成测试 │ ├── test_smart_scraper_integration.py │ ├── test_multi_graph_integration.py │ └── test_file_formats_integration.py ├── graphs/ # 图级测试(如 smart_scraper_openai_test.py) ├── nodes/ # 节点级测试(如 fetch_node_test.py) └── utils/ # 工具函数测试(如 convert_to_md_test.py)下文将沿"配置 → 基础设施 → 测试套件 → CI/CD"的顺序逐层深入。
二、核心配置:pytest.ini 逐项解读
仓库根目录的 pytest.ini 是整个测试体系的"总开关",它同时承载了测试发现、标记注册、覆盖率与超时等配置。
2.1 测试发现与收集规则
python_files = test_*.py *_test.py python_classes = Test* python_functions = test_* testpaths = tests- 只要文件名以
test_开头或以_test结尾都会被收集,因此test_smart_scraper_multi_concat_graph.py与smart_scraper_clod_test.py这类命名都能被识别; - 类名必须以
Test开头、函数名必须以test_开头; testpaths = tests限定默认只在 tests/ 目录内收集,配合norecursedirs排除.git、__pycache__、.venv、node_modules等目录,避免误收集第三方依赖。
2.2 默认运行参数(addopts)
addopts = -v # 详细输出 --tb=short # 精简回溯 --strict-markers # 未注册的 marker 直接报错 --cov=scrapegraphai # 覆盖率统计对象 --cov-report=term-missing --cov-report=html:htmlcov --cov-report=xml --cov-branch # 分支覆盖率 --durations=10 # 展示最慢的 10 个测试 -W default --strict-config --color=yes这些默认参数意味着直接运行pytest就会自动附带覆盖率统计与耗时排名,无需额外传参。
2.3 测试标记(Markers)
markers = integration: Integration tests requiring network access slow: Slow-running tests llm_provider: Tests for specific LLM providers requires_api_key: Tests requiring API keys benchmark: Performance benchmark tests unit: Unit tests (fast, no external dependencies) e2e: End-to-end tests共注册 7 个自定义标记。--strict-markers保证任何测试代码中出现的 marker 都必须在此处或 tests/conftest.py 中预先注册,从根源上杜绝拼写错误导致的静默跳过。
2.4 超时与异步配置
timeout = 300 asyncio_mode = auto全局单测超时为 300 秒(对应 pytest-timeout 插件),单个测试也可用@pytest.mark.timeout(120)覆盖。asyncio_mode = auto允许异步测试函数被 pytest-asyncio 自动收集执行。
2.5 覆盖率细节
[coverage:run]指定统计scrapegraphai包本体,并排除测试代码与site-packages;[coverage:report]通过exclude_lines忽略pragma: no cover、raise NotImplementedError、if TYPE_CHECKING:等"非业务"行;precision = 2将覆盖率精确到小数点后两位。
注意:
minversion = 8.0表示要求pytest 8.x及以上版本,运行前请确认环境中pytest --version满足要求。
三、共享 Fixtures 层:tests/conftest.py
tests/conftest.py 是整个测试体系的核心枢纽,它聚合了 LLM 提供商配置、Mock 模型、测试数据、临时文件、性能跟踪与自定义插件钩子。运行任何测试时,pytest 会自动加载该文件,无需手动 import。
3.1 多 LLM 提供商配置 fixtures
原文档强调"多提供商支持",conftest 中为每个支持的服务商都提供了即插即用的配置 fixture:
| Fixture | 默认模型 | 底层依赖 |
|---|---|---|
openai_config | gpt-3.5-turbo | OPENAI_APIKEY环境变量,缺省回落为test-key |
openai_gpt4_config | gpt-4 | 同上 |
ollama_config | ollama/llama3.2 | OLLAMA_BASE_URL,默认http://localhost:11434 |
anthropic_config | anthropic/claude-3-sonnet | ANTHROPIC_APIKEY |
groq_config | groq/llama3-8b-8192 | GROQ_APIKEY |
azure_config | azure_openai/gpt-35-turbo | AZURE_OPENAI_KEY+AZURE_OPENAI_ENDPOINT |
gemini_config | gemini/gemini-pro | GEMINI_APIKEY |
multi_llm_config | 参数化组合 | 依次展开 openai / ollama / anthropic / groq |
以 openai_config 为例,其返回结构完全对齐 ScrapeGraphAI 的图配置规范:
@pytest.fixture def openai_config() -> Dict[str, Any]: api_key = os.getenv("OPENAI_APIKEY", "test-key") return { "llm": { "api_key": api_key, "model": "gpt-3.5-turbo", "temperature": 0, }, "verbose": False, "headless": True, }temperature: 0保证输出确定性,headless: True关闭浏览器渲染,这些默认值都是为了"测试可重复"。multi_llm_config是参数化 fixture,用它在同一个测试函数上可以自动生成针对多个提供商的用例。
3.2 Mock LLM 与 Mock Embedder
单元测试不触碰真实 API,靠的是mock_llm_model与mock_embedder_model(tests/conftest.py):
@pytest.fixture def mock_llm_model(): mock = Mock() mock.model_name = "mock-model" mock.predict = Mock(return_value="Mocked LLM response") mock.invoke = Mock(return_value="Mocked LLM response") return mockMock 对象同时实现了predict与invoke两个接口,覆盖了不同版本 LangChain LLM 的调用约定,让底层节点(如 generate_answer_node.py)无需任何网络请求即可完成全链路单元测试。
3.3 测试数据与临时文件 fixtures
- 内容型 fixture:
sample_html(含产品/项目列表的标准 HTML)、sample_json_data(公司信息 JSON)、sample_xml(员工 XML)、sample_csv(人员 CSV),为不同格式的 ScraperGraph 提供输入; - 文件型 fixture:
temp_json_file、temp_html_file、temp_xml_file、temp_csv_file基于 pytest 内置的tmp_path将上述内容落盘为真实临时文件,返回文件路径字符串,正好匹配JSONScraperGraph、XMLScraperGraph、CSVScraperGraph等图对source参数的支持。
3.4 性能跟踪 fixtures
benchmark_config提供基准配置(warmup_runs: 1、test_runs: 3、timeout: 60);performance_tracker返回收集执行时间、Token 用量与 API 调用次数的指标字典,供测试自行记录。
3.5 Mock Server fixtures
mock_server在 localhost:8888 启动 HTTP 服务(详见下一节),mock_server_url给出其基础 URL,mock_website_url则指向可被TEST_WEBSITE_URL覆盖的线上测试站点。
3.6 自定义 pytest 钩子:按标记自动跳过
conftest 中定义了三个关键的钩子(tests/conftest.py):
pytest_configure:动态注册 integration / slow / llm_provider / requires_api_key / benchmark 五个标记,与 pytest.ini 中的注册互补;pytest_addoption:新增--integration、--slow、--benchmark三个开关型命令行参数;pytest_collection_modifyitems:在收集阶段自动给未加开关的 integration / slow 测试打上skip标记;对requires_api_key测试,则检查OPENAI_APIKEY、ANTHROPIC_APIKEY、GROQ_APIKEY三者中是否存在任一环境变量,若都没有就跳过——这让没有密钥的开发者也能安全地跑完整个测试目录。
四、Mock HTTP 服务器:tests/fixtures/mock_server
原文档重点介绍了一个"功能完整的本地 HTTP 服务器",其实现位于 tests/fixtures/mock_server/server.py。它基于标准库http.server与线程实现,无需安装任何外部服务,核心价值是让基础测试完全不依赖外网。
4.1 端点清单(与文档一致,含底层行为说明)
| 端点 | 返回内容 | 底层实现要点 |
|---|---|---|
/ | 首页(含导航链接) | 静态 HTML |
/products | 3 个产品条目(名称/描述/价格/库存状态) | 静态 HTML,带data-id、product-name等语义化 class |
/projects | 2 个项目条目(标题/描述) | 静态 HTML |
/api/data.json | 公司 + 员工 JSON | Content-type: application/json |
/api/data.xml | 公司 + 员工 XML | Content-type: application/xml |
/api/data.csv | 员工 CSV | Content-type: text/csv |
/slow | 延迟 2 秒后返回简单 HTML | 内部time.sleep(2) |
/error/404 | 404 错误页 | 真实 HTTP 404 状态码 |
/error/500 | 500 错误页 | 真实 HTTP 500 状态码 |
/rate-limited | 前 5 次正常返回,第 6 次起返回 429 +Retry-After: 60 | 按客户端 IP 计数(见下方代码) |
/dynamic | 带时间戳的动态内容 | 每次请求内容不同 |
/pagination?page=N | 每页 10 条、共 50 条的翻页列表 | 解析 query 参数生成 |
/rate-limited的实现揭示了它的工作方式(tests/fixtures/mock_server/server.py):按client_address[0]计数,超过 5 次后返回 429 并携带Retry-After: 60头——这正是测试重试逻辑与限流处理的理想靶场。
4.2 服务器生命周期
MockHTTPServer封装了完整生命周期:start()在后台守护线程中运行serve_forever;stop()执行shutdown()与server_close();get_url(path)拼接完整 URL;同时实现了上下文管理器(__enter__/__exit__),支持with MockHTTPServer() as server:的写法。conftest 中的mock_serverfixture 正是用 yield 模式确保测试结束后调用stop()清理资源。
典型用法:
def test_with_mock_server(mock_server): url = mock_server.get_url("/products") scraper = SmartScraperGraph( prompt="Extract products", source=url, config=config, )五、性能基准框架:tests/fixtures/benchmarking.py
benchmarking.py 是一套自研的轻量基准框架,围绕三个数据类与一个可装饰函数展开。
5.1 核心组件
BenchmarkResult(dataclass):单次运行的记录,字段包括test_name、execution_time、memory_usage(可选)、token_usage(可选)、api_calls、success、error、metadata;BenchmarkSummary:多次运行的统计摘要,包含num_runs、mean_time、median_time、std_dev、min_time、max_time、success_rate、total_tokens、total_api_calls;BenchmarkTracker:结果收集器,核心方法为:record(result):追加一条结果;get_summary(test_name):用statistics.mean/median/stdev计算该测试的统计量;save_results(filename="benchmark_results.json"):将全部结果导出为 JSON(自动创建benchmark_results/目录);generate_report():生成人类可读的文本报告,按测试名分组展示各项指标。
5.2 可复用的 benchmark() 函数
benchmark(func, name=None, warmup_runs=1, test_runs=3, tracker=None)以函数为单位执行"预热 + 多次实测"流程:预热轮异常被吞掉不计入;实测轮用time.perf_counter()计时并捕获异常写入BenchmarkResult;若被基准函数返回 dict 且含metadata键,会自动提取进结果。
5.3 基准 fixture 与基线对比
benchmark_trackerfixture(benchmarking.py)在测试结束后自动调用save_results()落盘,因此任何使用该 fixture 的基准测试都会自动产出 JSON 结果。
基线对比由pytest_benchmark_compare(baseline_file, current_file)承担:它加载基线 JSON 与当前 JSON,按test_name对齐,计算执行时间百分比变化;以 ±10% 为阈值,超过 +10% 记为回归(regressions),低于 -10% 记为改进(improvements),仅在当前出现的记为new_tests,返回结构化对比字典。实践中可按如下流程操作:
# 保存基线 pytest --benchmark -m benchmark cp benchmark_results/benchmark_results.json baseline.json # 运行新版本并对比 pytest --benchmark -m benchmarkfrom pathlib import Path from tests.fixtures.benchmarking import pytest_benchmark_compare comparison = pytest_benchmark_compare( Path("baseline.json"), Path("benchmark_results/benchmark_results.json"), )六、测试工具集:tests/fixtures/helpers.py
helpers.py 提供了可复用的断言、Mock 构造、数据生成与通用工具,用于提升测试的可读性与一致性。
6.1 断言辅助函数
assert_valid_scrape_result(result, expected_keys=None):断言结果非空且为 dict/str,可选校验必含键;assert_execution_info_valid(exec_info):校验图执行元信息(get_execution_info()的返回值)为合法 dict;assert_response_time_acceptable(execution_time, max_time=30.0):性能断言,默认上限 30 秒;assert_no_errors_in_result(result):在结果文本中检索error、exception、failed、timeout、rate limit等错误指示词,命中即失败。
6.2 Mock 响应构造
create_mock_llm_response(content, **kwargs):生成带content、response_metadata属性的 Mock 响应;create_mock_graph_result(answer, exec_info, error):返回(state, exec_info)二元组,state 内含answer/error键,模拟AbstractGraph的节点状态流转。
6.3 数据生成器
generate_test_html(title="Test Page", num_items=3, item_template="Item {n}"):定制标题与条目数生成 HTML;generate_test_json(num_records=3):生成含items(id/name/description/value)与total的 JSON;generate_test_csv(num_rows=3):生成id,name,value三列的 CSV。
6.4 校验与通用工具
validate_schema_match(data, schema_class):尝试用 Pydantic schema 类实例化数据,成功返回 True——与集成测试中传入schema=参数的图行为呼应;validate_extracted_fields(result, required_fields, min_values=1):校验提取字段存在且 list 长度达标;load_test_fixture(fixture_name)与save_test_output(content, filename):读写测试文件;compare_results(result1, result2, ignore_keys=None):忽略指定键后比较两个抓取结果;fuzzy_match_strings(str1, str2, threshold=0.8):基于分词集合重叠度的相似性判断(源码注释明确说明生产环境可改用 difflib/fuzzywuzzy);RateLimitHelper(max_requests, time_window):滑动窗口限流模拟器,can_make_request()/record_request()配合使用;retry_with_backoff(func, max_retries=3, initial_delay=1.0, backoff_factor=2.0):指数退避重试,延迟按 1s、2s、4s 递增,全部失败后抛出最后一次异常。
七、三层集成测试套件
原文档把集成测试划分为三个文件,位于 tests/integration/,全部以真实 LLM 连接验证端到端行为。
7.1 智能抓取集成(test_smart_scraper_integration.py)
tests/integration/test_smart_scraper_integration.py 面向SmartScraperGraph,覆盖场景包括:
- 多提供商/多场景:用
mock_server抓取/projects、/products,校验结果与执行信息; - Schema 化抓取:定义
ProjectSchema/ProjectListSchema(继承pydantic.BaseModel),传入schema=参数,验证结构化输出; - 超时处理:抓取
/slow端点,通过config["loader_kwargs"] = {"timeout": 5000}配置 5 秒加载超时; - 错误条件:抓取
/error/404,断言要么优雅返回、要么异常信息包含404/not found; - 真实网站:
TestRealWebsiteIntegration使用mock_website_url(默认指向测试站点,可由TEST_WEBSITE_URL覆盖)做真实网络抓取; - 性能基准:
TestSmartScraperPerformance用benchmark_tracker记录smart_scraper_basic的执行时间,并断言 < 30 秒。
典型集成测试骨架(直接来自源码):
@pytest.mark.integration @pytest.mark.requires_api_key class TestSmartScraperIntegration: def test_scrape_with_openai(self, openai_config, mock_server): url = mock_server.get_url("/projects") scraper = SmartScraperGraph( prompt="List all projects with their descriptions", source=url, config=openai_config, ) result = scraper.run() assert_valid_scrape_result(result) exec_info = scraper.get_execution_info() assert_execution_info_valid(exec_info)7.2 多图并发集成(test_multi_graph_integration.py)
tests/integration/test_multi_graph_integration.py 覆盖SmartScraperMultiGraph的并发抓取、多页面性能基准与SearchGraph集成,专门验证多源并行场景的正确性。
7.3 文件格式集成(test_file_formats_integration.py)
tests/integration/test_file_formats_integration.py 同时覆盖 JSON、XML、CSV 三种格式,每种都分别验证本地文件输入与Mock URL 输入两条路径,例如JSONScraperGraph既抓temp_json_file,也抓/api/data.json。
八、CI/CD 自动化流水线
原文档描述的 6 个 Job 在 .github/workflows/test-suite.yml 中均有对应实现,这里对照源码逐项说明。
8.1 触发器
on: push: branches: [main, pre/beta, dev] pull_request: branches: [main, pre/beta] workflow_dispatch:推送至 main / pre/beta / dev、PR 合入 main / pre/beta、以及手动触发三种方式都会启动流水线。所有 Job 都先执行uv sync安装依赖并playwright install chromium安装抓取所需的浏览器内核。
8.2 六个 Job 的职责与实现
| Job | 关键配置 | 要点 |
|---|---|---|
| unit-tests | 矩阵:3 个 OS × 3 个 Python 版本 | 命令pytest tests/ -m "unit or not integration" --cov --cov-report=xml;仅 Ubuntu+3.11 组合上传覆盖率到 Codecov |
| integration-tests | 矩阵:smart-scraper / multi-graph / file-formats | 注入OPENAI_APIKEY、ANTHROPIC_APIKEY、GROQ_APIKEY三个 secret,运行pytest tests/integration/ -m integration --integration -v;无论成败都上传htmlcov/与benchmark_results/产物 |
| benchmark-tests | 单 Job | 注入OPENAI_APIKEY,运行pytest tests/ -m benchmark --benchmark -v,上传benchmark_results/;PR 场景预留基线对比步骤 |
| code-quality | 单 Job | 依次执行ruff check scrapegraphai/ tests/、black --check、isort --check-only、mypy scrapegraphai/(mypy 为continue-on-error,不阻断流水线) |
| test-coverage-report | needs: [unit-tests, integration-tests] | 下载各 Job 覆盖率产物,PR 时通过python-coverage-comment-action在 PR 上评论覆盖率变化 |
| test-summary | needs: [unit-tests, integration-tests, code-quality] | 汇总输出各 Job 的最终状态 |
值得注意的工程细节:
fail-fast: false保证某个 OS/Python 组合失败不会取消其余组合;- 集成测试对真实 LLM 提供商的覆盖依赖 GitHub Secrets 注入密钥,本地无密钥时对应用例会被 conftest 自动跳过;
- 基准对比与覆盖率评论步骤在流水线中以"预留占位"方式存在,实际执行时会输出对应提示信息。
8.3 查看结果
- 单元测试与集成测试产物以 artifact 形式上传,可在 Actions 页面下载
htmlcov/与benchmark_results/; - 覆盖率 XML 通过 Codecov Action 上报;
- 基准 JSON 保留在
benchmark_results/中用于后续对比。
九、测试标记矩阵与灵活执行命令
9.1 标记一览
| Marker | 语义 | 典型用途 |
|---|---|---|
@pytest.mark.unit | 快速单元测试,无外部依赖 | Mock LLM 驱动的节点/工具测试 |
@pytest.mark.integration | 需网络访问的集成测试 | 真实 LLM + Mock Server / 真实网站 |
@pytest.mark.slow | 耗时长的测试 | 如/slow端点、真实网站抓取 |
@pytest.mark.benchmark | 性能基准测试 | 记录执行时间并断言上限 |
@pytest.mark.requires_api_key | 需要 API 凭据 | 无密钥时被自动跳过 |
@pytest.mark.llm_provider(name) | 指定 LLM 提供商 | 按提供商精确过滤 |
@pytest.mark.e2e | 端到端测试 | 完整抓取链路验证 |
多标记可叠加,例如 test_smart_scraper_integration.py 中真实网站测试同时挂了integration、slow、requires_api_key三个标记。
9.2 常用执行命令
# 全部测试(自动附带覆盖率与最慢 10 项统计) pytest # 仅单元测试(默认命令的等价写法) pytest -m "unit or not integration" # 集成测试(显式开启开关,否则会被自动跳过) pytest --integration # 性能基准 pytest --benchmark -m benchmark # 慢速测试 pytest --slow # 带 HTML 覆盖率报告 pytest --cov=scrapegraphai --cov-report=html # 运行指定文件 pytest tests/integration/test_smart_scraper_integration.py # 详细输出 pytest -v9.3 环境变量
集成测试依赖以下环境变量(对应 conftest 与工作流中的注入):
# LLM API 密钥 export OPENAI_APIKEY="sk-..." export ANTHROPIC_APIKEY="sk-ant-..." export GROQ_APIKEY="gsk_..." export GEMINI_APIKEY="..." # Azure OpenAI export AZURE_OPENAI_KEY="..." export AZURE_OPENAI_ENDPOINT="https://..." # 测试目标与本地模型 export TEST_WEBSITE_URL="https://scrapegrah-ai-website-for-tests.onrender.com" export OLLAMA_BASE_URL="http://localhost:11434"十、编写新测试:三类模板
原文档提供了三类可直接套用的模板,此处与仓库实际测试风格对齐后给出。
10.1 单元测试模板
import pytest from unittest.mock import Mock, patch class TestMyFeature: @pytest.fixture def setup(self): """Setup fixture for tests.""" return {"data": "value"} def test_my_function(self, setup, mock_llm_model): """Test description.""" # Arrange # Act # Assert10.2 集成测试模板
import pytest from scrapegraphai.graphs import SmartScraperGraph @pytest.mark.integration @pytest.mark.requires_api_key class TestMyIntegration: def test_real_scraping(self, openai_config, mock_server): """Test with real LLM provider.""" url = mock_server.get_url("/test-page") scraper = SmartScraperGraph( prompt="Extract data", source=url, config=openai_config, ) result = scraper.run() assert result is not None assert isinstance(result, dict)10.3 基准测试模板
import pytest import time from tests.fixtures.benchmarking import BenchmarkResult @pytest.mark.benchmark class TestMyBenchmark: def test_performance(self, benchmark_tracker, openai_config): """Benchmark test description.""" start = time.perf_counter() # Run operation to benchmark end = time.perf_counter() result = BenchmarkResult( test_name="my_benchmark", execution_time=end - start, success=True, ) benchmark_tracker.record(result)使用
benchmark_trackerfixture 后,测试结束会自动调用tracker.save_results(),将结果写入benchmark_results/benchmark_results.json(该目录由 BenchmarkTracker 自动创建)。
十一、常见问题排障
- 测试超时:全局超时在 pytest.ini 中为 300 秒;单测可用
@pytest.mark.timeout(120)单独放大。 - API 限流:优先改用 Mock Server 的
/rate-limited端点;需要程序化控制时使用RateLimitHelper(max_requests=5, time_window=60)。 - 测试不稳定(Flaky):可借助 pytest-rerunfailures 的
@pytest.mark.flaky(reruns=3, reruns_delay=2)自动重试。 - 无密钥导致跳过:集成测试被跳过属预期行为,配置第三节中的任一 API 密钥环境变量即可放行。
十二、当前路线与后续演进
原文档在 "Next Steps" 中给出了明确的演进方向,这些同样可作为社区贡献入口:
- 为更多图类型(如
DepthSearchGraph、SpeechGraph、ScreenshotScraperGraph等)补充集成测试; - 扩展 Mock Server 以覆盖更真实的场景(如表单提交、iframe、认证页);
- 引入基于截图对比的视觉回归测试;
- 通过 mutation testing 评估测试用例质量;
- 引入 Hypothesis 进行基于属性的测试;
- 建设性能趋势可视化看板;
- 针对并发抓取场景增加负载测试。
十三、给贡献者的测试约定
原文档对新增测试提出了明确规范,汇总如下:
- 优先复用 tests/conftest.py 中的既有 fixtures,避免重复造轮子;
- 为测试添加合适的标记(
@pytest.mark.*),确保过滤与自动跳过逻辑正确; - 遵循现有测试目录结构(
tests/graphs/、tests/nodes/、tests/utils/、tests/integration/)与命名约定(test_*.py/*_test.py); - 在文档中说明测试对 API 密钥、网络等外部条件的要求;
- 提交前确保本机
pytest全绿,CI 中的六类 Job 全部通过。
完整的运行与编写指引还可参考 tests/README_TESTING.md,它是这套基础设施配套的开发者手册,其中包含目录结构图、fixtures 清单与排障说明。
结语
ScrapeGraphAI 的增强测试基础设施,从 pytest.ini 的标记与覆盖率体系,到 tests/conftest.py 的多提供商 fixtures 与自动跳过钩子,再到 tests/fixtures/mock_server/server.py 的本地 HTTP 靶场与 benchmarking.py 的回归检测,最后汇入 .github/workflows/test-suite.yml 的跨平台 CI 流水线,构成了一条"无密钥也能跑、有密钥更全面、性能有监控、回归有预警"的完整质量保障链路。无论是为仓库补充新图类型的测试,还是在本机快速验证抓取逻辑,这套体系都能直接为你所用。
- 网页爬虫
- 人工智能
- AI 应用
【免费下载链接】Scrapegraph-ai
Python scraper based on AI
相关推荐
ClickHouse v24.10.2.80-stable 版本更新全解析:并行副本、查询优化与稳定性修复
ClickHouse v24.10.2.80 stable 版本更新全解析:并行副本、查询优化与稳定性修复 导读 本文围绕 ClickHouse® 实时分析数据
网页爬虫人工智能AI 应用QEMU 测试基础设施完全指南:从 make check 到容器化 CI 的测试体系详解
QEMU 测试基础设施完全指南:从 make check 到容器化 CI 的测试体系详解 导读 本文基于 QEMU 官方开发文档( docs/devel/tes
虚拟化硬件仿真pytest-mock配置完全指南:从基础到高级Mock测试优化
pytest mock配置完全指南:从基础到高级Mock测试优化 你是否在使用pytest mock时遇到过Mock版本冲突?是否因断言错误信息模糊而浪费数小时
开发工具
创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考