Hive Fidelity AI · Independent laboratory

Evaluation science for systems that can act.

We study the behavior, safety, and security of AI agents in the environments where they actually operate—not just the benchmarks where they look good.

What fidelity means here

Model scores are observations. Evaluation science asks whether they are true, attributable, and useful.

Agents are sociotechnical systems: the model is only one causal layer. Tools, prompts, memory, permissions, network conditions, environments, and verifiers can turn the same underlying capability into very different outcomes.

We design and interrogate evaluations that preserve those distinctions. The aim is not a prettier leaderboard. It is evidence strong enough to guide safety work, engineering decisions, and scientific disagreement.

FUTURE CONCEPT · BRIDGET 01Future concept rendering of Bridget, a bumblebee-inspired social support robot on a stable two-wheel base

Blue Ox Robotics · Original mission

Embodied intelligence must earn the right to be close to people.

We intend to help design, build, and rigorously evaluate assistive embodied AI for autistic children and adults—then, with a real safety case behind us, for people living with memory loss and dementia.

Bridget was our first design: a bumblebee-shaped social-support robot built around communication sovereignty, consent, physical safety, and companionship without coercion.

Explore the embodied-AI mission

Research domains

Where we work

Independent, evidence-led research across evaluation, safety, cyber, and agentic systems.

01

Agentic evaluation

We evaluate the system that acts: model, harness, tools, memory, environment, and verifier. Capability claims are only useful when the causal layer is clear.

02

Safety evaluation

Independent, adversarial evaluation of failure modes, misuse pathways, oversight assumptions, and safeguards—grounded in observable behavior and reproducible evidence.

03

Cybersecurity

Controlled research on cyber-capable agents, defensive workflows, environment integrity, and the boundary between benchmark performance and operational risk.

04

Evaluation research

Task validity, verifier science, contamination controls, uncertainty, calibration, and methods for learning what an evaluation result actually means.

Evaluation doctrine

Receipts before rhetoric.

Our methods are designed to survive contact with the result somebody hoped to get.

01

Outcome validity

A trajectory is not a success. We require the requested state to exist and the verifier to measure the right thing.

02

System attribution

We separate model capability from agent, prompt, tools, scaffolding, environment, and grading behavior.

03

Adversarial scrutiny

We test evaluations themselves for reward hacking, false confidence, hidden invalidity, and convenient stories.

04

Reproducible receipts

Claims travel with configurations, artifacts, limitations, and enough provenance for another researcher to challenge them.

Field program · Open source

The Usage-Surplus OSS Program

Subscription coding plans reset their usage windows whether we used the capacity or not. We turn the surplus into public maintenance.

Agents are dispatched to reproduce neglected bugs, repair evaluation infrastructure, update benchmark adapters, and do the unglamorous work that active open-source projects still need.

Surplus usage → tested upstream fixes.Agent authorship is disclosed. Human review gates maintainer-ready submissions.
Raccoon field operative carrying code patches and a wrench
Special Agent, Upstream Affairs

Public research notes

Blackwell Shenanigans

View all notes

Collaboration policy

Not for hire.
Still very much in the world.

Hive Fidelity is not accepting commercial engagements, consulting work, or paid client projects.

Open

Open-source work

Unpaid contributions where evaluation infrastructure, tooling, or scientific review can improve a public project.

Open

Third-party safety evaluation

Volunteer, independent evaluation of safety-relevant systems when scope, access, and responsible publication are appropriate.

Open

Research collaboration

Non-paid collaboration with researchers and builders on agent evaluation, cyber, safety, and measurement problems.

Good fits are public-interest, methodologically serious, and compatible with independent reporting.

Find our public work