Agentic evaluation
We evaluate the system that acts: model, harness, tools, memory, environment, and verifier. Capability claims are only useful when the causal layer is clear.
Hive Fidelity AI · Independent laboratory
We study the behavior, safety, and security of AI agents in the environments where they actually operate—not just the benchmarks where they look good.
What fidelity means here
Agents are sociotechnical systems: the model is only one causal layer. Tools, prompts, memory, permissions, network conditions, environments, and verifiers can turn the same underlying capability into very different outcomes.
We design and interrogate evaluations that preserve those distinctions. The aim is not a prettier leaderboard. It is evidence strong enough to guide safety work, engineering decisions, and scientific disagreement.

Blue Ox Robotics · Original mission
We intend to help design, build, and rigorously evaluate assistive embodied AI for autistic children and adults—then, with a real safety case behind us, for people living with memory loss and dementia.
Bridget was our first design: a bumblebee-shaped social-support robot built around communication sovereignty, consent, physical safety, and companionship without coercion.
Explore the embodied-AI missionResearch domains
Independent, evidence-led research across evaluation, safety, cyber, and agentic systems.
We evaluate the system that acts: model, harness, tools, memory, environment, and verifier. Capability claims are only useful when the causal layer is clear.
Independent, adversarial evaluation of failure modes, misuse pathways, oversight assumptions, and safeguards—grounded in observable behavior and reproducible evidence.
Controlled research on cyber-capable agents, defensive workflows, environment integrity, and the boundary between benchmark performance and operational risk.
Task validity, verifier science, contamination controls, uncertainty, calibration, and methods for learning what an evaluation result actually means.
Evaluation doctrine
Our methods are designed to survive contact with the result somebody hoped to get.
A trajectory is not a success. We require the requested state to exist and the verifier to measure the right thing.
We separate model capability from agent, prompt, tools, scaffolding, environment, and grading behavior.
We test evaluations themselves for reward hacking, false confidence, hidden invalidity, and convenient stories.
Claims travel with configurations, artifacts, limitations, and enough provenance for another researcher to challenge them.
Field program · Open source
Subscription coding plans reset their usage windows whether we used the capacity or not. We turn the surplus into public maintenance.
Agents are dispatched to reproduce neglected bugs, repair evaluation infrastructure, update benchmark adapters, and do the unglamorous work that active open-source projects still need.

Public research notes
Eighty-eight completely fabricated Hopper-adjacent sayings by Kirsten Ruge, drafted with Codex, Claude Fable 5.1, and GLM-5.2—historically disclaimed and spiritually plausible.
Google shipped Gemma 4 assistant drafters, SGLang merged Frozen-KV MTP support, and a single B300 turned it into a real 1.6x decode-speed win.
NVIDIA dropped a 31B activated-multimodal model with video, audio, image, OCR, GUI, tool-calling, and long-context support. That sounds suspiciously close to a pair-programming onboarding primitive.
This week’s frontier-model-in-a-small-Blackwell-shaped-box experiment ended with a useful answer: yes, Kimi K2.6 can fit, but only if you stop acting like the box is an H200.
Collaboration policy
Hive Fidelity is not accepting commercial engagements, consulting work, or paid client projects.
Unpaid contributions where evaluation infrastructure, tooling, or scientific review can improve a public project.
Volunteer, independent evaluation of safety-relevant systems when scope, access, and responsible publication are appropriate.
Non-paid collaboration with researchers and builders on agent evaluation, cyber, safety, and measurement problems.
Good fits are public-interest, methodologically serious, and compatible with independent reporting.
Find our public work