FAHALWAYS / EVALUATION RESEARCH
Study the judgment behind an AI result.
Commission research into how AI agents choose methods, interpret evidence, and complete technical work. We develop focused study proposals for research teams, model developers, and organizations investigating consequential failure modes.
Discuss a research scopeSelected research directions
The topics below describe proposed studies. They are not claims of completed agent experiments, customer deployments, or funded partnerships.
Choosing a valid evaluation
Does an agent select a test protocol that matches the real operating decision? A study can examine data leakage, unsuitable train/test splits, confounding, and whether limitations are disclosed.
Keeping a data pipeline correct
Can an agent choose appropriate record keys and revision rules when public feeds change format or restate earlier values? Examine whether a running pipeline preserves the meaning of the data.
Reproducing a reported number
Can an agent reconcile source versions and document assumptions when recreating an official statistic? Examine whether unresolved differences are explained rather than concealed.
What a commissioned study can include
- A research question, hypotheses, and an agreed analysis plan.
- Rights-cleared datasets, task packages, and reference materials.
- Specified model runs or experimental conditions and a reproducible record of the tested setup.
- Expert review criteria, error analysis, and uncertainty reporting.
- A technical report, reproducibility materials, and a findings discussion.
The final scope defines models, sample sizes, evaluation conditions, execution responsibilities, review process, schedule, budget, and permitted publication. No particular result is promised.
Public overview, private methods
We share the research question, intended contribution, and proposed deliverables during an initial discussion. Detailed task construction, reference materials, and scoring methods are disclosed through an agreed private review. Existing preprints can be discussed where relevant; publication and reuse rights are agreed separately.
Choose the right starting point
A three-case evaluation pilot develops a focused package for an agreed workflow. A research engagement investigates a broader question across a defined experimental design. We will help establish which scope matches your objective.
Discuss a research scopeExplore the evaluation pilot