still folding · tools

Still folding.

Two initial tools are on PyPI, and the leaderboard and benchmarks are coming soon.

Evaluate generated regular expressions: semantic equivalence, correctness, and ReDoS safety.

python apache-2.0 pypi 0.4.0

labloop

Inspired by autoresearch, an agent-driven experiment loop: propose a change, run it time-boxed, keep it only if the metric improves.

python apache-2.0 pypi 0.2.0

Runs regexbench across models and publishes the numbers: scores, methodology, and a re-run command.

evaluation planning

Reach out:

info@plicara.ai
← Back to the lab Benchmarks GitHub