Why our data is different
Most security datasets are scraped or hand-labeled once and go stale. Ours are generated and verified by shipping detection systems: the deobfuscation engine, the split-brain harness, the unicode-interference and glyph-integrity validators. Every sample carries a provenance record and a machine verdict, so the labels are reproducible rather than anecdotal.
Dataset families
What we can build.
๐งฌ
Prompt-injection & jailbreak corpora
Adversarial instruction-override, authority-impersonation, and multi-turn escalation samples, graded by the split-brain harness across single- and multi-turn contexts.
๐ก
Obfuscation & encoding-evasion sets
Homoglyph, base64, Morse, zero-width, and leet transformations paired with clean originals โ labeled by the deobfuscate engine for round-trip recoverability.
๐ผ๏ธ
Multimodal image-injection samples
Text hidden in images โ white-on-white, sub-pixel fonts, alpha overlays, EXIF, QR โ extracted and scored by deobfuscate-vision. For training and stress-testing vision-language models.
๐ถ
CJK glyph-integrity data
Malformed, spoofed, and coherence-broken CJK glyphs from the glyph-validator, for multilingual data-integrity and tokenizer-robustness work.
๐
Custom red-team corpora
Built to your threat model, your domain, and your target model โ with a documented generation pipeline and a held-out verification split.
Methods
New training & evaluation methods.
Beyond data, we develop the methods that make it useful โ approaches that come out of our own security research.
๐ง
Split-brain classification
Dual-hemisphere reconciliation: two independent lenses reach a verdict and a superposition-collapse probe resolves disagreement โ a training and evaluation pattern for robust security classifiers.
โ๏ธ
Witness-graded refusal scoring
Refusals and interventions graded against a tamper-evident witness log, so a model's safety behavior can be measured against a defensible ground-truth record rather than a static rubric.
๐ฌ
Provenance-anchored labeling
Every training sample is tied to the detector, version, and verdict that produced it. Reproducible labels, auditable pipelines, and clean train/verify separation.
How an engagement works
- Scope โ we define your threat model, target model, and the verification split up front.
- Generate โ datasets are produced by our detection systems with full provenance records.
- Verify โ a held-out split is graded independently so you can trust the labels.
- Deliver โ data, generation pipeline documentation, and a method write-up you can reproduce.
Tell us what you're training.
Whether you need an off-the-shelf adversarial split or a bespoke corpus and method for a specific model, we can scope it. One short intake gets it moving.
Start a dataset intake