Code Story episode

S11 Bonus: The 77% Token Tax: Why Fixed-Token Chunking is Poisoning Your Enterprise RAG Pipelines and How to Fix the Ingestion Layer with Dr. Alex Kihm, Founder & CEO of POMA AI

Aug 28, 2025 · Season 11 · 25 min

Solving hallucinations in LLMs... through the chunks.

Alex Kihm just turned 40, and has been into computers for 36 years. He was given his first hand me down computer at the age of 4 by his parents, and also grew up with lego bricks and building things. He is an engineer by training, but eventually switched to econometrics on the big data side. Outside of his professional life, he is married and describes himself as water affectionate. He enjoys swimming, diving - and free diving. In fact, he studied diving during his semester abroad. His free diving is mainly a hobby, but he has deep respect for the professional free divers of the world.

At his original startup, Alex started to dive into LLMs and immediately ran into RAG (Retrieval Augmented Generation). When he observed the wrong information being returned, along with a ton of resource consumption in the process - IE cost - he set out to solve the problem, and figured out the solution was in the chunks.

This is the creation story of POMA AI.

Sponsors

Paddle.comSema SoftwarePropelAuthPostmanMeilisearchMailtrap.TECH Domains (https://get.tech/codestory)Links

https://poma-ai.com/https://www.linkedin.com/in/alex-kihm-27a902338/

Checkout our episode stacks on Stacklist! https://stacks.codestory.co/

Advertising Inquiries: https://redcircle.com/brands

Privacy & Opt-Out: https://redcircle.com/privacy

Full show notes and links on codestory.co

More episodes

Keep listening

Topics

Explore related conversations