Michael Mabinuola

Database/AI Engineer | Performance, Vector Search and Data Pipelines

Core skills

  • Distributed database architecture
  • OLAP and OLTP
  • Query performance tuning
  • Linux performance
  • High availability
  • S3 object storage
  • Vector search and RAG
  • Data pipelines and ETL
  • Python · SQL · Bash
  • SingleStore · MySQL · ClickHouse

About

I keep databases fast and make data searchable by meaning.

Data has to be stored safely and come back out in milliseconds, ideally with the right answer. That’s my work, from database performance and vector search down to the systems underneath. And if there’s another layer under that, I’m going in.

Latest

What I’m reading and learning right now.

Experience

Where I’ve worked and what I did there.

Database Engineer · Agile Platform (에이플랫폼) Dec 2023 – present

  • Diagnose and resolve SingleStore performance bottlenecks for on-premise customers using query profiles, execution plans and OS-level metrics.
  • Tune SingleStore for vector data storage and retrieval; benchmark and evaluate vector index performance.
  • Build semantic search systems and RAG-based AI pipelines, and provide technical support for AI search and vector functions.

Data Engineer / Data Analyst · Starpickers (별따러가자) Sep 2021 – Nov 2023

  • Designed and implemented end-to-end ETL pipelines processing 1M+ rows daily from 300+ users.

Education: MBA, Chung-Ang University · BSc (Honors) Economics, Don State Technical University

Languages: English (native) · Russian (proficient) · Korean (intermediate)

Vector search

I build search that understands meaning, not just matching words.

  • 1024-dim IVF_FLAT (baseline): 38 GB index memory, 11.3 QPS, recall@5 97.8%
  • 128-dim IVF_FLAT, 100 candidates: 5 GB index memory, 74.9 QPS, recall@5 96.0%
  • 64-dim IVF_FLAT, 100 candidates: 2.52 GB index memory, 114.9 QPS, recall@5 87.8%

Real work: Vector search with Matryoshka embeddings

Articles

I write up what I test, so others can repeat it.

  • Vector search with Matryoshka embeddings · SingleStore blog, 2026 [EN]
    Two-stage vector search on 10M vectors: 87% less index memory and 6.6x throughput at 96% recall. I ran the benchmarks; co-authored with SingleStore engineers.
  • High-dimension vectors in SingleStore · Agile Platform wiki, 2025 [KO]
    Tested indexing, inserts, and KNN/ANN search on vectors from 16K to 131K dimensions (up to 98 GB), and showed the real insert limit is max_allowed_packet, not the column type.
  • Hamming distance for SingleStore · Agile Platform wiki, 2026 [KO]
    A Rust WebAssembly UDF that adds Hamming distance over binary vectors, which SingleStore lacks natively.
  • On-premises RAG with Llama 2 and SingleStore · Agile Platform wiki, 2026 [KO]
    A CPU-only RAG example: a quantized Llama 2 model answers from context retrieved out of a SingleStore vector store, using LangChain and llama.cpp.

More write-ups exist but are internal to my employer.

Database analysis

I find out why a database is slow and fix it.

Real work: Query profile analysis tool (in progress) · Unlimited storage PoC on VAST object storage (Validated VAST as S3-compatible storage for SingleStore: lifecycle tests (backup, HA, PITR, scaling), TPC-H benchmarks on hot and cold tiers, and throughput under load.)

Open source

When a tool breaks, I find out why and tell the people who can fix it.

Root-caused a startup failure on Oracle Linux 10

  • Symptom: All four agent sessions failed within a second of starting the daemon.
  • Cause: Oracle Linux 10 ships tmux as a pre-release snapshot (reports next-3.4). Its server crashes on capture-pane. Reproduced with plain tmux and an empty config, so OpenRig and my config were ruled out.
  • Fix: Built stable tmux 3.5a from source, then recovered the sessions.
  • Outcome: Maintainer merged the findings into the getting-started guide (PR #999).

OpenRig · bug report · Read the issue on GitHub

Side projects

Older things I built for fun, because I like making pages move.

From before AI wrote the CSS. Unrelated to my day job.

  • BITwallet: A landing page design for a crypto payments app. (HTML, CSS)
  • SAYN: A photography site design concept, inspired by Rron Berisha. (HTML, CSS)
  • Simple survey: A survey app built with Streamlit. (Python, Streamlit)

Contact

Email is best. I reply during Korean working hours (KST, UTC+9).

Measured for real

The database looked like it had memory to spare. Linux was quietly doing the caching.

On an on-premise SingleStore host with 7.5 GB RAM, I dropped the Linux page cache and ran two queries against a 33,554,432-row columnstore table.

  • select count(*) grew the page cache by only 9 MiB (answered from metadata).
  • select sum(g), sum(id) took 4.20 s and grew the page cache by 66 MiB.

On-premise columnstore data is cached by the Linux page cache through buffered reads, so it never shows up in the database engine’s own memory metrics.

Sizes are illustrative; counters and behaviour mirror cgroup v2.