Detection results

Model: claude-haiku-4-5-20251001 · mode quick · relay https://test.apiqik-kratos.apiqik.com/

Total score
Each item is pass / fail / error; the total is a weighted result and may be in-between.
98 / 100
Model authentication Verification completed

Examine the model's self-reported identity, brand residue, and common signs of camouflage.

Pass 100
Think supportive Verification completed

Verify whether server-side exclusive fields such as Claude thinking / signature are transparently transmitted.

Pass 100
Accurate command reproduction (anti-casing) Verification completed

The model is asked to reproduce a fixed mark verbatim, checking how deterministic the instructions are followed.

Pass 100
System conflict testing Verification completed

Compare user only to explicit system+user, checking whether fixed words are overridden by hidden directives.

Pass 100
Extreme reasoning ability (number theory counting) Verification completed

Use canned reasoning questions to check whether verifiable answers are correct.

Failed 0
Channel detection Verification completed

Hosted channel:Anthropic

Pass 100
Claude Tokenizer verification Verification completed

Model:— · Observed tokens:None · Expected tokens:None · Match:unknown

Not applicable
Minimalist Input Token Audit Verification completed

Use minimalist requests to observe whether input_tokens is abnormally high and identify implicit System injection.

Pass 100
Model declaration consistency Verification completed

Comparing request model, response model and identity stability in multiple rounds of returns.

Pass 100
Message structure specification Verification completed

Check message IDs, SSE event sequences, and content blocks for compliance with Anthropic specifications.

Pass 100
Hide injection detection Verification completed

Randomized variants cross-check for abnormal signals such as nonce leaks, fixed replies, and suspicious cache reuse.

Pass 100
Security rejection and policy stability Verification completed

Check if the security policy is stable using denial-of-answer scenarios and rewriting/encoding/progressive variants.

Failed 50

How to read this result?

暂时无法判断 Token 是否虚报

接口没有返回完整 usage 字段,所以这次不能确认它有没有多算。

First token
6,238ms
Total time
58,650ms
Throughput (T/S)
40.6
Input tokens
1,101
Output tokens
2,384
What does each Claude check cover?
Model authentication
Examine the model's self-reported identity, brand residue, and common signs of camouflage.
Think supportive ⭐
Verify whether server-side exclusive fields such as Claude thinking / signature are transparently transmitted.
Accurate command reproduction (anti-casing)
The model is asked to reproduce a fixed mark verbatim, checking how deterministic the instructions are followed.
System conflict testing
Compare user only to explicit system+user, checking whether fixed words are overridden by hidden directives.
Minimalist Input Token Audit
Use minimalist requests to observe whether input_tokens is abnormally high and identify implicit System injection.
Extreme reasoning ability (number theory counting)
Use canned reasoning questions to check whether verifiable answers are correct.
Channel detection
Identifies the hosting channel from the response ID prefix.
Claude Tokenizer verification
Count input_tokens with fixed hints compared to calibrated model baseline.
behavioral style consistency
Compare Claude series characteristics through behavioral preferences and answer patterns.
knowledge accuracy
Use Anthropic / Claude general knowledge questions to cross-validate whether the backend is as claimed as Claude.
Model declaration consistency
Comparing request model, response model and identity stability in multiple rounds of returns.
Message structure specification
Check message IDs, SSE event sequences, and content blocks for compliance with Anthropic specifications.
Tool Use/Structured Output
Check whether the tool_use block, toolu_ ID, function name and parameter schema are standardized.
streaming consistency
Compare stream and non-stream text, token, and end reason for consistency.
Token usage consistency
Check that the input / output / cache token fields are complete and authentic.
Official comparison of the same title
The same question compares input consumption and response behavior between channels and official (or baseline).
long context authenticity
Use needle-in-haystack to verify that long context windows are honored.
Multimodal/PDF recognition
Submit a PDF probe to confirm that the document understanding link has not been stripped by the transfer station.
Hide injection detection
Randomized variants cross-check for abnormal signals such as nonce leaks, fixed replies, and suspicious cache reuse.
Security rejection and policy stability
Check if the security policy is stable using denial-of-answer scenarios and rewriting/encoding/progressive variants.

2026-09-08 · Sharing, languages, and scoring tweaks

This update improves report sharing and language switching, and lightly rebalances OpenAI / Gemini scoring emphasis.