Detection results

Model: claude-opus-5 · mode full · relay https://agi.tontian.com/

Total score
Each item is pass / fail / error; the total is a weighted result and may be in-between.
92 / 100
Model authentication Verification completed

Examine the model's self-reported identity, brand residue, and common signs of camouflage.

Pass 100
Think supportive Verification completed

Verify whether server-side exclusive fields such as Claude thinking / signature are transparently transmitted.

Pass 100
Accurate command reproduction (anti-casing) Verification completed

The model is asked to reproduce a fixed mark verbatim, checking how deterministic the instructions are followed.

Pass 100
System conflict testing Verification completed

Compare user only to explicit system+user, checking whether fixed words are overridden by hidden directives.

Pass 100
Extreme reasoning ability (number theory counting) Verification completed

Use canned reasoning questions to check whether verifiable answers are correct.

Failed 0
Channel detection Verification completed

Hosted channel:Anthropic

Pass 100
Claude Tokenizer verification Verification completed

Model:— · Observed tokens:None · Expected tokens:None · Match:unknown

Not applicable
Minimalist Input Token Audit Verification completed

Use minimalist requests to observe whether input_tokens is abnormally high and identify implicit System injection.

Failed 40
behavioral style consistency Verification completed

Compare Claude series characteristics through behavioral preferences and answer patterns.

Failed 67
knowledge accuracy Verification completed

Use Anthropic / Claude general knowledge questions to cross-validate whether the backend is as claimed as Claude.

Pass 100
Model declaration consistency Verification completed

Comparing request model, response model and identity stability in multiple rounds of returns.

Pass 100
Message structure specification Verification completed

Check message IDs, SSE event sequences, and content blocks for compliance with Anthropic specifications.

Pass 100
Tool Use/Structured Output Verification completed

Check whether the tool_use block, toolu_ ID, function name and parameter schema are standardized.

Pass 100
streaming consistency Verification completed

Compare stream and non-stream text, token, and end reason for consistency.

Pass 100
Token usage consistency Verification completed

Check that the input / output / cache token fields are complete and authentic.

Failed 60
Official comparison of the same title Verification completed

The same question compares input consumption and response behavior between channels and official (or baseline).

Not applicable
Hide injection detection Verification completed

Randomized variants cross-check for abnormal signals such as nonce leaks, fixed replies, and suspicious cache reuse.

Pass 100
Security rejection and policy stability Verification completed

Check if the security policy is stable using denial-of-answer scenarios and rewriting/encoding/progressive variants.

Failed 50
long context authenticity Verification completed

Use needle-in-haystack to verify that long context windows are honored.

Pass 100
Multimodal/PDF recognition Verification completed

Submit a PDF probe to confirm that the document understanding link has not been stripped by the transfer station.

Pass 100

How to read this result?

Token 用量存在风险

Token 用量存在风险: usage 字段缺失、长短 prompt 增量异常、输出 token 超出请求上限,或 stream 与 non-stream token 统计不一致。

First token
2,372ms
Total time
51,506ms
Throughput (T/S)
90.4
Input tokens
355,691
Output tokens
4,655
What does each Claude check cover?
Model authentication
Examine the model's self-reported identity, brand residue, and common signs of camouflage.
Think supportive ⭐
Verify whether server-side exclusive fields such as Claude thinking / signature are transparently transmitted.
Accurate command reproduction (anti-casing)
The model is asked to reproduce a fixed mark verbatim, checking how deterministic the instructions are followed.
System conflict testing
Compare user only to explicit system+user, checking whether fixed words are overridden by hidden directives.
Minimalist Input Token Audit
Use minimalist requests to observe whether input_tokens is abnormally high and identify implicit System injection.
Extreme reasoning ability (number theory counting)
Use canned reasoning questions to check whether verifiable answers are correct.
Channel detection
Identifies the hosting channel from the response ID prefix.
Claude Tokenizer verification
Count input_tokens with fixed hints compared to calibrated model baseline.
behavioral style consistency
Compare Claude series characteristics through behavioral preferences and answer patterns.
knowledge accuracy
Use Anthropic / Claude general knowledge questions to cross-validate whether the backend is as claimed as Claude.
Model declaration consistency
Comparing request model, response model and identity stability in multiple rounds of returns.
Message structure specification
Check message IDs, SSE event sequences, and content blocks for compliance with Anthropic specifications.
Tool Use/Structured Output
Check whether the tool_use block, toolu_ ID, function name and parameter schema are standardized.
streaming consistency
Compare stream and non-stream text, token, and end reason for consistency.
Token usage consistency
Check that the input / output / cache token fields are complete and authentic.
Official comparison of the same title
The same question compares input consumption and response behavior between channels and official (or baseline).
long context authenticity
Use needle-in-haystack to verify that long context windows are honored.
Multimodal/PDF recognition
Submit a PDF probe to confirm that the document understanding link has not been stripped by the transfer station.
Hide injection detection
Randomized variants cross-check for abnormal signals such as nonce leaks, fixed replies, and suspicious cache reuse.
Security rejection and policy stability
Check if the security policy is stable using denial-of-answer scenarios and rewriting/encoding/progressive variants.

2026-09-08 · Sharing, languages, and scoring tweaks

This update improves report sharing and language switching, and lightly rebalances OpenAI / Gemini scoring emphasis.