Adding Tools & Integrations — Definition of Done
Use this checklist whenever you add or materially change:- a tool — under
integrations/<vendor>/tools/for a single-vendor tool, ortools/system//tools/cross_vendor/for a cross-cutting one (see tool-placement-policy.md) - an integration under
integrations/<name>/— its config, client, verifier, and tools
1. Tool checklist
Files usually involved
integrations/<vendor>/tools/<tool_name>_tool/__init__.py— the tool package (most common path: the tool belongs to a vendor integration)tools/system/<tool_name>/ortools/cross_vendor/<tool_name>/— only when the tool is not vendor-specific (e.g.tools/system/sre_guidance_tool/)integrations/<name>/client.py— reuse a dedicated integration API client instead of inlining requestscore/tool_framework/utils/— shared helper code reused across vendorsdocs/<tool_name>.mdx— user-facing usage, parameters, examplestests/tools/test_<tool_name>.py— behavior and regression coverage
tools/ and integrations/<vendor>/tools/, so placement is about ownership, not discovery — see tool-placement-policy.md. Wherever a tool lives, it calls integration-local clients/helpers rather than inlining transport, and never lives in a top-level vendors/ or services/ package.
Tool packages must be substantive production modules — no empty or discovery-only __init__.py, no thin wrapper that only satisfies registry import. Any tool with validation, credential/parameter resolution, transport/client calls, output normalization, or error handling should split those concerns into focused sibling files (tool.py, models.py, validation.py, delivery.py/client.py, results.py), leaving __init__.py as a small registry entrypoint that imports the public tool object.
Contract and implementation
- Pick the simplest shape that fits (
@tool(...)for lightweight tools, a richer class only when needed) -
__init__.pyis a small registry entrypoint; non-trivial tools use sibling modules for implementation concerns - Metadata is complete and accurate:
name,description,source,surfaces,requires, and anyuse_cases/outputs/retrieval_controls -
input_schemamatches the actual runtime arguments and required fields -
is_availablereturnsTrueonly when the tool can genuinely run -
extract_paramsmaps resolved integration state into tool args correctly - Validation, credential/parameter resolution, transport/client calls, and result formatting are separated so each can be tested independently
- Reusable transport or integration-specific parsing lives in
integrations/<name>/orcore/tool_framework/utils/, not copied into the tool body - Failure responses have a stable, agent-friendly shape; expected external failures (missing config, auth, rate limit, upstream 4xx/5xx) return structured errors rather than raising — unexpected exceptions use the global
BaseToolwrapper intentionally or are migrated with telemetry coverage - Output is normalized enough for the agent/LLM to consume reliably
- Secrets never leak through
extract_params, return values, logs, or traceable tool-call kwargs; secret/PII output is run throughinfrastructure/safety/masking/before return - External side effects declare
side_effect_level,requires_approval, andapproval_reasonwhere appropriate - To appear in both chat and action turns, set
surfaces=(ToolSurface.CHAT, ToolSurface.ACTION)
Live payload parsing
If the tool parses API, MCP, log, or webhook payloads:- Validate against the real or documented upstream response shape, not only idealized mocks
- Handle alternate field names used in live payloads
- Handle missing or partial fields without returning unusable output
- Preserve important context when truncating, tailing, paginating, or flattening data
- Upstream 429 / 5xx responses return a clear, agent-friendly error rather than raising
- Add at least one regression test using a realistic fixture payload
hasMore / cursor mismatches; content-vs-pointer shapes (logs_content vs logs_url-style payloads).
Skill guidance (optional)
A tool can carry workflow guidance the model reads on every call by shipping aSKILL.md. The guidance is appended to the tool’s description under a Workflow guidance: heading — it is permanent schema text, not a side channel. All guidance targeting one tool is combined and truncated at 2400 characters (tools/registry_skill_guidance.py), so budget it like description text: the longer the guidance, the more of every request it consumes.
Skill guidance and a harness playbook are independent and composable — a tool may have both. Skill guidance rewrites one tool’s description with call rules (parameters, refusals, what the tool owns); a harness playbook under core/agent_harness/prompts/skills/ describes the multi-step workflow around it. The GitHub tools (github_cli, ci_fix, security_fix) each ship a SKILL.md; ci_fix and security_fix also have a workflow card. Give the two different names (operating-github-ci-fixer vs fixing-github-ci) and keep tool-level facts in the tool card only. Placement rules: core/agent_harness/prompts/skills/AGENTS.md.
When to add (two independent axes — evaluate both):
Neither is the default for most of ~70
tools/ packages. Missing SKILL.md is usually correct. Never add one-per-vendor stubs, or one skill per observability vendor when the failure mode is shared (query hygiene is one class, not Datadog + Grafana + CloudWatch copies).
Harness authoring template: core/agent_harness/prompts/skills/_template/SKILL_TEMPLATE.md.
File. A SKILL.md with YAML frontmatter and a markdown body:
- Explicit: add the file’s path to
_skill_guidance_files()intools/registry_skill_guidance.py. ASKILL.mdthat exists on disk but is absent from that tuple is never loaded — unlisted and missing files are skipped with no diagnostic, the same trap as forgetting adocs.jsonentry. - Or place it under
tools/system/python_execution_tool/skills/*/SKILL.md, which is discovered automatically.
unknown_tool— a name undertools:matches no registered tool; that target is dropped.invalid_metadata— missingname/description/tools, an over-longname/description, or anamethat is not lowercase kebab-case; the whole skill is skipped.parse_failed— malformed frontmatter YAML.
2. Integration checklist
Files usually involved
integrations/<name>/__init__.py— package facade: a docstring, plus re-exports of the public API when callers need them (see the__init__.pyrule in AGENTS.md)integrations/<name>/config.py— config model,classify(), validators, selectors, normalization helpersintegrations/<name>/client.py— a dedicated API client, when the integration makes direct remote callsintegrations/<name>/verifier.py— local verification logicintegrations/<name>/tools/<tool_name>_tool/— the vendor’s agent-callable tools (see §1)integrations/catalog.py— resolve the integration into the shared runtime configintegrations/verify.py— wire the local verification pathdocs/<name>.mdx— user-facing setup, usage, verificationtests/integrations/test_<name>.py, plustests/tools/ortests/e2e/where tools or scenarios exercise it
integrations/<name>/ owns everything about one vendor — config, resolution, clients, verifiers, helpers, and its tools. Only vendor-less (tools/system/) and cross-vendor (tools/cross_vendor/) tools live under top-level tools/.
Examples from the repo
- Datadog:
integrations/datadog/(withintegrations/datadog/tools/),integrations/catalog.py, tests undertests/integrations/datadog/andtests/tools/test_datadog_*.py. - Grafana:
integrations/grafana/(withintegrations/grafana/tools/),integrations/catalog.py,surfaces/cli/wizard/local_grafana_stack/, tests undertests/integrations/grafana/andtests/tools/test_grafana_*.py. - Bitbucket:
integrations/bitbucket/shows the facade layout — a one-line__init__.pybesideconfig.py,client.py,verifier.py, andtools/.
Core completeness
- Config, normalization, validators, and
classify()are in place underintegrations/<name>/config.py, leaving__init__.pya facade - Catalog resolution / env loading is wired correctly
- Verification path is wired in
integrations/verify.pyand adapters/registry as needed - Integration-local client added under
integrations/<name>/client.py(only if it makes direct remote calls) - Tool layer is wired and stable
- CLI setup flow is updated if the integration is user-configurable locally
-
opensre integrations setup <name>parity is added, or intentionally documented as out of scope - New required env vars / credentials are added to
.env.example(never.env) - Sensitive credentials follow the Credential resolution contract below
-
make verify-integrationspasses
Credential resolution
Secret env names and non-secret config follow different write/read paths. Keep this contract when adding or changing an integration.
Hard rules for new code
- Never use bare
os.getenvfor a secret env name (*_TOKEN,*_KEY,*_PASSWORD,*_SECRET, connection strings, and similar). Useresolve_env_credentialfromconfig.llm_credentials(env first, then the credentials file). - Webhook /
*_URLvalues are never written to the credentials file (wizard routes them to store/.env, notsync_env_secret). Read with store → plainos.getenvonly. Webhook URLs often embed a secret token — treat them like passwords for logging/masking. - Leave
load_env_integration_servicesplain-env-only (startup-safe; no credentials-file read at boot). - Store still wins in
resolve_effective/ merge — env/credentials-file is the fallback tier only. - Tools receive credentials through
extract_params(resolved integration state), never their own env reads. At execution, keys listed in the tool’sinjected_paramsoverride model-supplied values, so the verified source wins even when the model passes a token. A tool’s resolver may read env only as the final fallback when nothing was injected —integrations/github/tools/github_cli/credentials.pyis the reference: explicit/injected token first, thenGITHUB_MCP_AUTH_TOKEN, thenGITHUB_TOKEN/GH_TOKEN. - Set
OPENSRE_DISABLE_KEYRING=1to skip local-file reads/writes (env and store still work).
resolve_env_credential (env → credentials file), sync_env_secret / save_credential (secret writes to the credentials file and .env), sync_env_values (.env keys, including secrets).
3. Discovery and edge cases
For tools that list, search, or inspect resources:- Folder/nested resource layouts are considered where the upstream supports them
- Large result sets are capped or paginated intentionally
- Partial fetches are surfaced clearly (
truncated,fetch_error, etc.) - Time/order-sensitive results preserve causal ordering where it matters
4. Docs and tests
Docs
- Ship or update a
docs/page/section in the same PR (new tool, CLI command, agent behavior, or integration; and whenever a tool’s API/schema or an integration’s setup changes) - Any new
docs/page is registered indocs/docs.json(without the.mdxsuffix) so Mintlify navigation shows it
Tests
- Unit tests for config/normalization
- Tool contract tests, or equivalent schema/metadata coverage
- A registry/discovery test proves the tool is visible on the expected surface(s)
- Runtime behavior tests for success and failure paths
- At least one realistic fixture for live-payload parsing when external payloads are involved
- A test proves the agent can discover or invoke the tool through the normal runtime path
-
tests/integrations/updated when integration wiring changes
Final gate (new integrations)
Everything above is complete, and:- Screenshot or demo GIF showing the integration working end-to-end
- E2E test added
- CI checks pass (see CI.md)