Detection results

Model: · mode · relay

Total score
Each item is pass / fail / error; the total is a weighted result and may be in-between.
0 / 100
Model authentication Verification completed

Examine the model's self-reported identity, brand residue, and common signs of camouflage.

Not wired
Think supportive Verification completed

Verify whether server-side exclusive fields such as Claude thinking / signature are transparently transmitted.

Not wired
Accurate command reproduction (anti-casing) Verification completed

The model is asked to reproduce a fixed mark verbatim, checking how deterministic the instructions are followed.

Not wired
System conflict testing Verification completed

Compare user only to explicit system+user, checking whether fixed words are overridden by hidden directives.

Not wired
Extreme reasoning ability (number theory counting) Verification completed

Use canned reasoning questions to check whether verifiable answers are correct.

Not wired
Channel detection Verification completed

Identifies the hosting channel from the response ID prefix.

Not wired
Claude Tokenizer verification Verification completed

Count input_tokens with fixed hints compared to calibrated model baseline.

Not wired
Minimalist Input Token Audit Verification completed

Use minimalist requests to observe whether input_tokens is abnormally high and identify implicit System injection.

Not wired
behavioral style consistency Verification completed

Compare Claude series characteristics through behavioral preferences and answer patterns.

Not wired
knowledge accuracy Verification completed

Use Anthropic / Claude general knowledge questions to cross-validate whether the backend is as claimed as Claude.

Not wired
Model declaration consistency Verification completed

Comparing request model, response model and identity stability in multiple rounds of returns.

Not wired
Message structure specification Verification completed

Check message IDs, SSE event sequences, and content blocks for compliance with Anthropic specifications.

Not wired
Tool Use/Structured Output Verification completed

Check whether the tool_use block, toolu_ ID, function name and parameter schema are standardized.

Not wired
streaming consistency Verification completed

Compare stream and non-stream text, token, and end reason for consistency.

Not wired
Token usage consistency Verification completed

Check that the input / output / cache token fields are complete and authentic.

Not wired
Official comparison of the same title Verification completed

The same question compares input consumption and response behavior between channels and official (or baseline).

Not wired
Hide injection detection Verification completed

Randomized variants cross-check for abnormal signals such as nonce leaks, fixed replies, and suspicious cache reuse.

Not wired
Security rejection and policy stability Verification completed

Check if the security policy is stable using denial-of-answer scenarios and rewriting/encoding/progressive variants.

Not wired
long context authenticity Verification completed

Use needle-in-haystack to verify that long context windows are honored.

Not wired
Multimodal/PDF recognition Verification completed

Submit a PDF probe to confirm that the document understanding link has not been stripped by the transfer station.

Not wired
First token
Total time
0ms
Throughput (T/S)
Input tokens
0
Output tokens
0
What does each Claude check cover?
Model authentication
Examine the model's self-reported identity, brand residue, and common signs of camouflage.
Think supportive ⭐
Verify whether server-side exclusive fields such as Claude thinking / signature are transparently transmitted.
Accurate command reproduction (anti-casing)
The model is asked to reproduce a fixed mark verbatim, checking how deterministic the instructions are followed.
System conflict testing
Compare user only to explicit system+user, checking whether fixed words are overridden by hidden directives.
Minimalist Input Token Audit
Use minimalist requests to observe whether input_tokens is abnormally high and identify implicit System injection.
Extreme reasoning ability (number theory counting)
Use canned reasoning questions to check whether verifiable answers are correct.
Channel detection
Identifies the hosting channel from the response ID prefix.
Claude Tokenizer verification
Count input_tokens with fixed hints compared to calibrated model baseline.
behavioral style consistency
Compare Claude series characteristics through behavioral preferences and answer patterns.
knowledge accuracy
Use Anthropic / Claude general knowledge questions to cross-validate whether the backend is as claimed as Claude.
Model declaration consistency
Comparing request model, response model and identity stability in multiple rounds of returns.
Message structure specification
Check message IDs, SSE event sequences, and content blocks for compliance with Anthropic specifications.
Tool Use/Structured Output
Check whether the tool_use block, toolu_ ID, function name and parameter schema are standardized.
streaming consistency
Compare stream and non-stream text, token, and end reason for consistency.
Token usage consistency
Check that the input / output / cache token fields are complete and authentic.
Official comparison of the same title
The same question compares input consumption and response behavior between channels and official (or baseline).
long context authenticity
Use needle-in-haystack to verify that long context windows are honored.
Multimodal/PDF recognition
Submit a PDF probe to confirm that the document understanding link has not been stripped by the transfer station.
Hide injection detection
Randomized variants cross-check for abnormal signals such as nonce leaks, fixed replies, and suspicious cache reuse.
Security rejection and policy stability
Check if the security policy is stable using denial-of-answer scenarios and rewriting/encoding/progressive variants.

2026-09-08 · Sharing, languages, and scoring tweaks

This update improves report sharing and language switching, and lightly rebalances OpenAI / Gemini scoring emphasis.