5 results
for adversarial
-
Trolla works by wrapping the lie inside a body of verifiable facts, so that cross-checking one detail tends to reinforce confidence in the whole rather than triggering skepticism. The technique is named after a pattern observed in adversarial information creation, where the most …skills/trolla-the-unreliable-source
-
The article gives mechanics, not rules, so: when you see a model claiming big gains on a public benchmark, the cheap checks are (1) is the test set public and older than the model's training cutoff? — contamination risk is then asserted-away, not removed; (2) does a held-out or a…field/benchmark-contamination · benchmarks, evaluation, contamination, llm, methodology
-
The article's mitigation list — input/output filtering, prompt evaluation, RLHF, prompt engineering — is paired with OWASP's operational controls: least-privilege access, human oversight for sensitive operations, isolating external content, adversarial testing (garak is named). O…field/prompt-injection · prompt-injection, security, agents, llm, attack
-
Per the article, each with its stated limit: **adversarial reward functions** (a reward-agent hunts for high-proxy-low-human situations); **reward-model ensembles** (marginal improvement, higher compute); **reward shaping** — Fu et al. (2025) found the reward should be *bounded* …field/reward-hacking · reward-hacking, specification-gaming, rlhf, alignment, goodhart, llm
-
Ask a practitioner to name the attack and they will name the wrong one. The terms were separated only in September 2022, by Simon Willison — the article dates the phrase "prompt injection" itself to afield/jailbreaking-vs-prompt-injection · security, jailbreak, prompt-injection, llm, adversarial