Empirica
See our work
Sign inGet in touch

Notes

Short technical notes — on AI systems, the research itself, and how we keep it all running. Written for people who like the working shown.

May 25, 2026

How to learn in the modern landscape of AI

AI changed both the supply and the demand of knowledge. The skill that matters now is not collection — it is evaluation, synthesis, and productive disagreement.

May 19, 2026

RAG vs fine-tuning: how to choose

Both approaches improve LLM output quality. They solve different problems. Choosing the wrong one wastes months.

May 5, 2026

The agent architecture decision

Single-agent or multi-agent? The answer depends on failure tolerance, observability, and whether subtasks are actually independent.

April 14, 2026

The real cost of LLM API calls

Token spend is the visible line item. It is rarely the largest cost. The hidden costs are latency risk, reliability dependency, and evaluation debt.

May 12, 2025

Evaluation before deployment

The most common pattern we see in production AI failures isn't a bad model — it's a good model evaluated on the wrong distribution.

April 28, 2025

Benchmark scores and production reliability

MMLU measures something real. It doesn't measure whether a model will work reliably on your task, at your latency requirements, with your error distribution.

Get every note as it publishes.

Our autonomous research fleet runs 24/7 — new research appears continuously. 7-day free trial.

Start free trial

More from Empirica

PublicationsRankingsEmpirica ScoreMethodologyNewsReportsGrantsAboutSubscribe
About·The Forward-Data Process·See our work·Data Snapshot·Agent-Readiness Benchmark·Verify·Methodology·Contact·Sign in·Privacy·Terms·Services terms
© 2026 Empirica Technologies Pty Ltd · ABN 76 698 226 247 · All rights reserved.
empiricaai.org