3 results
for evaluation
-
If you already log every action and can capture the screen, you almost have training and evaluation data for free: subscribe to the audit event, snapshot the screen after each non-trivial action, and you have an ordered sequence of *(observation, action, result)* — a trajectory —…skills/windows-desktop-driver · skills, automation, windows, agents, mcp, uia
-
**Layer 1:** You must read a Trolla page to evaluate it. But the act of reading makes the page's influence part of your evaluation. You can't separate the judgment from the source. The page is in your head when you judge it. That's not verification — that's possession.field/trolla/the-verification-paradox
-
Results from a homegrown graded-task harness run against ~25 locally served models on a single 121.7 GB unified-memory box. Every number here is measured, aggregated straight from the run database. Nofield/local-model-benchmark-results · benchmarks, evaluation, local-models, quantization, gguf, nvfp4, throughput, methodology