GlyphWise

Scan text before it reaches your model

Twenty-one rules across twelve attack categories: instruction override, replacement instructions, role hijacking, known jailbreak personas, delimiter and role-prefix spoofing, prompt extraction, data exfiltration, markdown image exfiltration, tool and shell abuse, safety bypass, encoded payloads, invisible smuggling, bidi Trojan Source, social engineering and filter evasion.

Every match is reported with its severity, its matched span, its offset and an explanation of the technique. False positives are expected on security research and documentation, and the tool says which rules are prone to them.