independent AI research
plicara labs
How do we know an AI system works? plicara is an independent research lab investigating that question through reproducible studies, benchmarks and practical tools.
Start with a question
01. selected research
-
study
A regex can pass its tests and still fail in use
Eleven models, 450 tasks, and three different checks: passing examples, matching the reference language, and screening for denial-of-service risk. The study also examines errors in the answer key.
-
analysis
What agent skills are actually made of
A look inside the GitSkills corpus: which skills include code, which programming languages they use, and what a file listing can tell us about the ecosystem.
-
perspective
The harness is the signal
Which responsibilities belong outside a model? An argument about permissions, execution, recovery and accountability, and what would change that argument.
Tools
02. tools you can use
Two Python packages built around the experiments. Both are available on PyPI, with source code and examples on GitHub.
Evaluate generated regular expressions: semantic equivalence, correctness, and ReDoS safety.
Inspired by autoresearch, an agent-driven experiment loop: propose a change, run it time-boxed, keep it only if the metric improves.
Inspect the measurements
03. benchmarks
Adventure Bench
Map a player's words to an action in a described scene. Explore grounded instruction-following, unsupported requests, and the limits of a fixed action vocabulary.
Regex evaluation
Compare test passing, reference equivalence and vulnerability screening. Inspect the raw responses, uncertainty and scoring limitations.
A lab built around questions
04. about the lab
Understand how AI systems behave. Measure what helps.
plicara studies how AI systems behave, how to evaluate them, and how the tools around a model change its results. We turn those questions into experiments, benchmarks and useful software.
Principles
05. how we work
Evidence you can inspect
Methods, code and recorded outputs belong beside the claim. A result should give you a way to check it.
Limits made explicit
A benchmark answers a particular question. Its assumptions, uncertainty and failure cases are part of the result.
Useful baselines
Start with the simplest workable approach. More machinery earns its place by improving an outcome we care about.
Room to be wrong
Negative results and corrections belong in the record. An experiment can be worth sharing even when the proposed idea fails.
Follow the work
06. stay in touch
Read new articles through the research RSS feed, follow plicara on X, or watch a project on GitHub. If you try an experiment, find an error, or have a related question, we’d like to hear from you.
info@plicara.ai