Community研究&データ分析github.com

1Lany2a/token-usage-report

Agent Skill: report per-conversation token usage, cache hit rate and estimated cost (DeepSeek peak/off-peak pricing) for ZCode, Codex, Claude Code, OpenCode

token-usage-report とは?

token-usage-report is a Claude Code agent skill that agent Skill: report per-conversation token usage, cache hit rate and estimated cost (DeepSeek peak/off-peak pricing) for ZCode, Codex, Claude Code, OpenCode.

対応Claude CodeCodex CLI~CursorOpenCode
npx skills add 1Lany2a/token-usage-report

Installed? Explore more 研究&データ分析 skills: obra/superpowers, affaan-m/quarkus-verification, affaan-m/uspto-database · View all 6 →

お気に入りのAIに質問する

このエージェントスキルを事前に読み込んだ状態で新しいチャットを開きます。

ドキュメント

Token Usage Report

Reads the agent's local rollout files and prints the usage of the most recent conversation segment plus the whole session (two lines):

本段会话:输入 123456 / 缓存命中 98765 / 输出 4321 / 命中率 80.0% / 费用 ¥0.12
整个会话:输入 987654 / 缓存命中 876543 / 输出 43210 / 命中率 88.8% / 费用 ¥0.50

Why both lines? The whole-session numbers are computed locally by the script (no extra model calls), so reporting them costs only a few dozen extra characters in the reply — negligible token overhead, and the user can compare the current turn against the session total.

When to run

  • Run before composing the final reply of every session.
  • Trigger phrasings: "token usage", "用量", "费用", "cost", "billing", "多少 tokens", "汇报", "spent".
  • Why: it gives the user a per-turn usage line they can cross-check against the provider's real bill. The cost is an estimate (pricing snapshot), not the bill.

How to run

Script: scripts/usage_report.py — Python 3.9+, standard library only (no install), except the optional zstandard package needed for DeepSeek Harness (dsh) logs.

python3 <this-skill-dir>/scripts/usage_report.py --latest   # most recent segment/session (default)
python3 <this-skill-dir>/scripts/usage_report.py --whole    # whole latest session
python3 <this-skill-dir>/scripts/usage_report.py --all      # per-session list
python3 <this-skill-dir>/scripts/usage_report.py --agent zcode|codex|claude|opencode|dsh
python3 <this-skill-dir>/scripts/usage_report.py --session <substring>   # pin a session

Always use the absolute path to this skill's directory. Do not cd elsewhere first. Pass --agent explicitly when you run it from inside an agent session (auto-detection picks the most recently modified file, which may belong to a different concurrently running agent); --session pins a specific session file.

Append the script's output to the final reply verbatim — both lines, on separate lines (本段会话 above, 整个会话 below). If the reply is markdown, a single newline renders as a space and merges the two lines, so put them inside a fenced code block (or leave a blank line between them):

```text
本段会话:输入 123456 / 缓存命中 98765 / 输出 4321 / 命中率 80.0% / 费用 ¥0.12
整个会话:输入 987654 / 缓存命中 876543 / 输出 43210 / 命中率 88.8% / 费用 ¥0.50
```
  • Exit code 0 → always include the lines, even if the numbers look identical to the previous report (the user wants them every turn).
  • Codex reports a single line (whole session) — there is no segment data.
  • Exit code 1 / "No usage data found" → skip silently; do not ask the user.
  • Any other error → fix the cause (missing Python, moved skill dir) before responding; do not silently skip.

Supported agents

AgentData source--latest means
ZCode~/.zcode/cli/rollout/model-io-*.jsonllast segment: records since the latest request whose last message is a user prompt (that turn incl. its tool calls)
Codex~/.codex/sessions/**/rollout-*.jsonllast session (whole) — Codex rollouts store cumulative running totals, so only the final total is reported, as a single line; segment data is unavailable
Claude~/.claude/projects/**/*.jsonllast segment: records since the latest "type":"user" record
OpenCode<data>/opencode/opencode.db (SQLite; %LOCALAPPDATA% on Windows, $XDG_DATA_HOME or ~/.local/share otherwise)last segment of the most recent session (from the latest user message)
DSH~/.dsh/sessions/**/session.jsonl[.zstd] (zstd-compressed; needs optional zstandard)last turn (one user exchange) + whole session

All four are parsed from documented, locally-verified or well-known formats. DeepSeek Harness (dsh) and other agents can be added with a small parser — see README.md for the recipe.

Accuracy & limitations

The report is an estimate from local rollout data + list prices, not the provider bill. Differences vs the official usage page can come from:

  • Codex totals are cumulative: each total_token_usage record is the running session total, so the script uses the last record (never sums). Per-request and per-segment numbers are unavailable for Codex.
  • Peak/off-peak is applied per session start time, while the provider prices each request by its own time. A session crossing a peak boundary (09:00–12:00 / 14:00–18:00 Beijing) is mispriced for the crossed part.
  • Reasoning output tokens are included in output for Codex (billed at the output rate); ZCode rollouts expose no reasoning field.
  • Local rollouts may not capture every request on the key (aborted streams, non-interactive calls), so totals can be lower than the full key usage.
  • Rates are a snapshot; a gateway/aggregator key or a different model ID may bill at different prices.

Cost model

DeepSeek official peak/off-peak pricing, effective 2026-08-17 00:00 Beijing time (announced 2026-08-13). Peak hours: 09:00–12:00 and 14:00–18:00 Beijing; all other times are off-peak at half price. Sessions before 2026-08-17 use the flat legacy rates.

ModelPeriodcached-inmiss-inout(CNY / 1M tokens)
deepseek-v4-flashoff-peak0.051.54.5
deepseek-v4-flashpeak0.103.09.0
deepseek-v4-prooff-peak0.154.513.5
deepseek-v4-propeak0.309.027.0

The billing period comes from the session start time (machine-local time). If the official rates change, update the PRICES_* tables in scripts/usage_report.py; the authoritative source is platform.deepseek.com.

Integration (report every session automatically)

Add this rule to the agent's user-level instruction file (~/.zcode/AGENTS.md, ~/.codex/AGENTS.md, CLAUDE.md, or a project AGENTS.md):

Run python3 <skill-dir>/scripts/usage_report.py --agent <your-agent> --latest (ALWAYS pass --agent — auto-detection can read another concurrently running agent's data) before every final reply and append its two output lines verbatim: 本段会话:输入 X / 缓存命中 Y / 输出 Z / 命中率 N% / 费用 ¥M 整个会话:输入 X / 缓存命中 Y / 输出 Z / 命中率 N% / 费用 ¥M. Skip only when the script prints "No usage data found".

Security

  • The script reads only local rollout JSON files under the user's home directory. It never reads, prints, or requires API keys, .env files, or credentials.
  • Rollout files may contain prompt text — never echo their contents. Print only the aggregated numbers the script outputs.

関連スキル