Arte Mentis

Independent AI consulting · San Diego

Recent Work

Numbers from real engagements, measured on real batches. No projections, and nothing a vendor deck would print without a footnote.

Case studies

≈45× tenant recall

Address Enrichment for a B2B Sales Team

A sales team’s address list was finding 8 businesses where hundreds existed, because commercial addresses hide many tenants behind one street number. I built an enrichment service that combines map data with premises-aware matching, and on the test batch it raised the count from 8 businesses to about 368: roughly 45 times the recall.

It now runs nightly in production with no manual involvement.

7 models tested

The Local Model Benchmark

I designed a benchmark that put seven language models (the small, local kind a business can run on its own hardware) through JSON extraction, mathematical reasoning, and executive summarization against fixed thresholds. Pass or fail, no partial credit. One model passed. I had built the thresholds hoping more would clear them.

Towards AI published the methodology and results in December 2025.

200+ trained

Training a Distributed Team

An AI evaluation firm needed more than 200 distributed professionals onboarded and producing consistent work. I built the training materials, task guidelines, and rubric-based playbooks; performance variance across the group dropped by 40 percent.

Want a number like these for your business?

Tell me what you’re trying to figure out, and I’ll tell you honestly whether I’m the right person for it.

Start a conversation
Macrocystis pyrifera, giant kelp. Off this coast it grows two feet a day. Results you can measure.