How does a text watermark work?
A first-principles investigation of how generation-time text watermarking works, from a weighted coin to Gemma.
Building production AI applications and foundational model systems with robust evaluation at scale. Writing about what ships.
A first-principles investigation of how generation-time text watermarking works, from a weighted coin to Gemma.
Testing if applying Matryoshka Representation Learning (MRL) to tabular entity data could bridge the cost gap, compressing embeddings enough to remain operational while still beating BM25 on corrupted queries.

A production diary of tiered episodic memory in AI agents. Three markdown files, an SQLite database, and a lobster that somehow remembers what you said three days ago.

Four tools, a settings file, and full control - how I built my own AI coding agent with Pi instead of paying $200/month for a CLI that keeps changing.

Practical way of using colbert with ragatouille on modal labs

Practical strategies for implementing and optimizing Retrieval Augmented Generation (RAG) in LLM systems.

Hackathon Experience at Mistral Hackathon in San Francisco



