Investment Data Platform
A large Singaporean investment holding company
Backend services and pipelines behind deal discovery, due diligence, and portfolio monitoring.
10+ interconnected microservices4 vendor data feeds processed
Data Engineer II at Thinking Machines Data Science — pipelines, agentic AI, MLOps, and cloud infrastructure for financial services, investment management, education, and compliance.

Tools I Work With
My work spans data pipelines, agentic AI, data warehousing, data modeling, MLOps, backend software engineering, and cloud infrastructure across multiple enterprise engagements.
I also speak about generative AI, AI coding agents, and modern engineering workflows through technical talks and community events.
A major Philippine bank
A daily pipeline that tells a bank who its ~15 million records actually are.
15M → 7M records consolidated to unique customer keys748,000 duplicate pairs surfaced by record linkage>99.9% five-tier confidence accuracy4h → <1h pipeline runtime
Builds production data and AI infrastructure for enterprise clients across financial services, investment management, education, and compliance.
Investment Data Platform
Enterprise Document Intelligence Platform
Built the backend of a compliance case-management platform on Django, deployed across two AWS regions with separate databases for data residency and GDPR. Decomposed the monolith into cross-region services and shipped ~50 production API endpoints supporting 100+ reports per day.
Built a Google Cloud data-processing pipeline for a COVID-19 contact-tracing and vaccination-tracking application and led a two-week data science internship for eight participants.
Co-authored and published the IEEE AIComprehend paper, then built and deployed the full-stack Django application used in a month-long controlled study that showed a 13.9% improvement in student test scores.
Modeled and transformed production datasets into Looker Studio dashboards used by the Technology and Operations team to track a 300-plus participant developer program.
01
Warehouses, lakehouses, and orchestration — moving and modeling data at enterprise scale.
02
Agentic systems, RAG, and probabilistic matching — prototype to production.
03
Containerized, infrastructure-as-code deployments across AWS, Azure, and GCP.
04
APIs and services behind enterprise platforms — FastAPI and Django at multi-region scale.
05
The core languages behind production data systems, pipelines, and services.
06
AI-assisted development and the collaboration stack that ships it.

2019 – 2024
BS Computer Engineering

2013 – 2019
High School Diploma, STEM
Drag a match-weight threshold and watch duplicate record clusters merge in real time — the same probabilistic linkage that rebuilt a bank's Single Customer View.

A proof-of-concept Model Context Protocol server that lets ChatGPT query and operate a Databricks workspace through a set of typed tools.

Led the organizing team behind 10+ in-person events over 2.5 years, bringing together 2,000-plus participants, ~25 speakers, ~25 partners and sponsors, and 50+ volunteers — growing the community from zero to 500-plus official members.
Member of the World Economic Forum-backed network of young leaders driving local impact through community projects.
Conference Talk
Configuring AI coding assistants with subagents, agent skills, and AGENTS.md.
Startup Community Event
Working with AI agents and modern engineering workflows.
Community Talk
Invited speaker on cloud and data engineering.
Conference · Doha, Qatar
Presented the AIComprehend research paper.
IEEE · Published Nov 27, 2023
Co-authored an adaptive reading-comprehension platform; a four-week controlled study with 58 high school students showed a 13.9% improvement in test scores.
Read the PaperUniversity of the Philippines Diliman · Jul 2024
Graduated summa cum laude in BS Computer Engineering with a 1.15 weighted average, finishing in the Top 5 of the program.
Apache Software Foundation · 2025
Contributed documentation covering Azure Blob Storage remote logging and Google Cloud Vertex AI operators.
Kyle took the messiest part of our data platform and turned it into the most dependable one. He communicates like a consultant and ships like a senior engineer.
The rare engineer who is equally at home in a C-suite steering meeting and a 2 a.m. pipeline incident. His record-linkage work changed how our bank sees its customers.
Every team Kyle touches gets faster. He automates the boring parts, documents the sharp edges, and leaves the codebase better than he found it.
Still the canonical map of the data-systems territory. The chapter on partitioning pays for the whole book.
GuideAnthropicThe clearest writing on when you need an agent and when a workflow is enough.
Specmodelcontextprotocol.ioThe protocol behind my ChatGPT-to-Databricks proof of concept. Read the spec before wrapping another API.
LibraryMoJ Analytical ServicesProbabilistic record linkage at scale. This is the engine behind the 748,000 duplicate pairs story.
Blogsimonwillison.netThe most consistently grounded coverage of LLMs and coding agents anywhere.
8,224 contributions in the last year