<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>Mathieu Ajaka — Writing</title><description>Technical notes on building and evaluating production AI systems — RAG, agents, evaluation, and time-series ML.</description><link>https://www.mathieuajaka.com/</link><language>en</language><item><title>Purged walk-forward vs ordinary cross-validation: preventing leakage in financial time-series ML</title><link>https://www.mathieuajaka.com/writing/purged-walk-forward-vs-cross-validation/</link><guid isPermaLink="true">https://www.mathieuajaka.com/writing/purged-walk-forward-vs-cross-validation/</guid><description>Why k-fold cross-validation leaks on overlapping-label time series, and the purging, embargo, CPCV, and deflated-Sharpe discipline that make evaluation honest.</description><pubDate>Wed, 29 Jul 2026 00:00:00 GMT</pubDate><category>time-series</category><category>backtesting</category><category>evaluation</category><category>leakage</category></item><item><title>Chunking Arabic legal text for RAG (and why BM25 collapses on clitics)</title><link>https://www.mathieuajaka.com/writing/chunking-arabic-legal-text-for-rag/</link><guid isPermaLink="true">https://www.mathieuajaka.com/writing/chunking-arabic-legal-text-for-rag/</guid><description>Why Arabic legal retrieval is hard, how naive BM25 silently fails on clitic-heavy queries, and the morphology-aware chunking and hybrid retrieval that fix it.</description><pubDate>Wed, 15 Jul 2026 00:00:00 GMT</pubDate><category>rag</category><category>arabic-nlp</category><category>retrieval</category><category>evaluation</category></item></channel></rss>