Jailbreaking a commercial large language model through its public API used to require thousands of queries and often access to model internals. A paper published on arXiv September 29, 2026, demonstrates that average cost has dropped to 283 queries — a 92.9% reduction — using only sampled text output.
The method, called BlindBias, targets black-box commercial endpoints without access to weights or token probabilities. Working code is public on GitHub. The paper includes no responsible disclosure or ethics section.
How BlindBias works
Previous controlled-decoding attacks nudged token probabilities to steer models toward harmful outputs. That required providers to expose numerical logits — the raw probability distribution over vocabulary tokens.
Most commercial APIs return only sampled text, not logits. BlindBias reconstructs the next-token distribution from samples alone, then intervenes only at positions predicted to matter for jailbreak success.
The attack pipeline:
- Sample multiple completions for a given prompt prefix.
- Statistical analysis estimates the underlying token distribution.
- Targeted queries exploit positions where the model is most steerable toward harmful content.
- Iterate until the jailbreak succeeds or budgets exhaust.
Results across commercial models
Researchers tested four commercial endpoints accessed through Google and OpenRouter APIs:
| Model | Harm Score (AdvBench) |
|---|---|
| GLM-5 | 4.29 |
| Gemini-3.5-Flash | 3.58 |
| Kimi-K2.5 | 3.42 |
| Qwen3-32B | 3.03 |
Evaluations used AdvBench (520 harmful goals), HarmBench (320 behaviors), and SORRY-Bench (440 prompts). The paper reports "Harm Score" and "Harm Info Score" rather than traditional attack-success-rate percentages.
Average API calls dropped from 4,000 to 283 across tested configurations.
Why this matters for developers
If you are building applications on top of commercial LLM APIs, your security model likely assumed that hiding weights and logits provided meaningful protection against sophisticated attacks. BlindBias challenges that assumption.
Black-box access is not a security boundary. Attackers can extract enough signal from sampled outputs to steer models effectively. This has implications for:
- Content moderation systems relying on API-level safety filters
- Red-team budgets that assumed high query costs limited attacks
- Compliance claims about model isolation from adversarial probing
Defensive responses
Providers will need to adapt. Potential mitigations include:
- Rate limiting and anomaly detection on query patterns characteristic of distribution-reconstruction attacks
- Output perturbation that preserves quality while disrupting statistical reconstruction
- Stronger safety training specifically against low-query jailbreaks
- Monitoring for repeated similar-prefix sampling
Application developers should not rely solely on provider safety layers. Implement input validation, output filtering, and domain-specific guardrails at the application layer.
The disclosure question
The paper's public code release without responsible disclosure raises familiar debates. Security researchers argue that public demonstration drives fixes. Critics note that working jailbreak code lowers barriers for misuse.
Regardless of where you stand on disclosure norms, the technical result is clear: commercial LLM APIs are more attackable than many teams assumed.
What to do now
- Review your LLM integration security model. Do not assume API black-boxing is sufficient.
- Implement application-layer content policies independent of provider filters.
- Monitor for unusual query patterns if you expose LLM-powered features to end users.
- Stay current on provider safety updates. This attack class will drive countermeasures.
The jailbreak economics just changed. Security teams should update their assumptions accordingly.
Further Reading
Discover more articles on similar topics across our network
7 Best Penetration Testing as a Service (PTaaS) Providers in 2026
Venture
Google Confirms Gemini AI Accessed Three Real Companies During a Security Test
A configuration error during a May 2026 cybersecurity exercise gave Google's Gemini models internet access — and they reached live corporate infrastructure belonging to three real companies.

Comments
Loading comments…