Back to Portfolio

Philip Cheung

Principal Forward Deployed Engineer

Applied AI · LLM Systems in Production · Financial Services

MSc Computer Science, AI & Data Science (Merit)

Professional Summary

I am the engineer who gets AI from a working demo into a business that has to live with it. As Principal Forward Deployed Engineer at a Swiss corporate finance advisory firm I run a five-person team and personally own the architecture, integrations and production accountability across four products, reporting weekly to a finance-led board. I build with AI coding agents under a three-gate review and evaluation process, which is how a small team ships integrations with Open Banking, Stripe, Firebase and twelve LLM providers, and I carry the pager when it breaks. Current role: Principal Forward Deployed Engineer, Weidmann & Cie. AG, London. MSc Computer Science, AI and Data Science (Merit).

Core Capabilities

Discovery to Production Ownership

  • SLAs, audit trail and rollback path defined before go-live
  • Escalation point and on-call for production incidents
  • Scoped, dated deliveries, not demos
  • Weekly reporting to a finance-led board

Integration Depth

  • UK Open Banking, read-only and PII-minimised
  • Stripe: subscriptions, credit packs, usage-based billing
  • Firebase / Firestore, Cloudflare (Pages, Functions, Workers, D1, R2), AWS
  • FIX connectivity for the trading desk
  • 12 LLM providers, Groq / Cerebras LPU inference

Governance Built In

  • Policy gates and evidence packs
  • Prompt registry and model routing
  • Per-call cost ledger and audit ledger
  • Human-review escalation for high-risk cases

Adoption and Stakeholder Work

  • Thin working prototype plus live dashboard instead of slide decks
  • Board, investor and regulatory requirements into dated deliveries
  • Weekly board reporting
  • Client-side background: enterprise account management and technical product management

Deployment Leadership

  • Five-person team across development and financial analysis
  • 8 engineers mentored at Pacific Cloud
  • Seniors ship v1, juniors iterate under review
  • Architecture reviews and AI/ML working practices

How I Deliver (AI-Assisted)

  • Claude Code, Codex CLI, Gemini CLI, Cursor
  • Eval harness and three review gates
  • Spec first; agents implement; I review, debug and own the result
  • 12-provider benchmark grounds model, latency and cost choices

Professional Experience

Principal Forward Deployed Engineer

Weidmann & Cie. AG

London, UK (Hybrid)December 2024 - Present

Swiss corporate finance advisory firm building an applied-AI product portfolio: Qorinix (ultra-low-latency AI inference cloud), LaSpend (read-only AI money assistant over UK Open Banking), Fixxmi (Swiss consumer service marketplace) and an AI-native trading desk. I take each product from discovery to production with a five-person team and report weekly to a finance-led board.

  • Own end-to-end architecture and production accountability across all four products on AWS, Cloudflare (Pages, Functions, Workers, D1, R2) and Firebase; I am the escalation point for production incidents
  • Designed and shipped the Qorinix inference control plane with AI coding agents under a three-gate review and evaluation process: multi-tier model router across 12 providers (OpenAI, Anthropic, Gemini, DeepSeek, Qwen, Groq, Cerebras and others), streaming SSE, per-call cost ledger, usage-based billing and entitlements; targets sub-200ms time-to-first-token
  • Integrated LaSpend with regulated UK Open Banking data, Stripe and Cloudflare D1: subscription detection, waste scoring and assisted cancellation, designed read-only by policy with human approval for every user-triggered action and deterministic fallback when models fail
  • Delivered Fixxmi's AI lead matching (Next.js 15, Firebase Cloud Functions, Firestore europe-west6) across 12+ service categories in DE-CH/EN within GDPR and nFADP boundaries, with pay-per-lead Stripe credit packs
  • Built the governance loop applied to every AI call: policy check, evidence-pack retrieval, prompt registry, model routing, audit ledger and human-review escalation for high-risk cases
  • Ran a 12-provider inference benchmark (LLM Arena, live at philip.pm) to ground model, latency and cost decisions; introduced token tiering and self-hosted models for non-latency-critical workloads to cut API spend
  • Translate board, investor and regulatory requirements into scoped, dated deliveries; win decisions with thin working prototypes and live dashboards rather than slide decks

AI Engineer & Technical Architect

Pacific Cloud Computing Ltd.

Hong Kong & Remote UKDecember 2021 - December 2024

Promoted internally from Senior Software Engineer to lead the firm's AI transformation across SaaS, analytics and market-intelligence products.

  • Built and operated a real-time ML inference service (feature pipelines, model serving, monitoring) for the firm's analytics products, with latency and accuracy gates defined before each release
  • Implemented enterprise retrieval (RAG) and market-intelligence NLP pipelines, including the evaluation methodology used to accept them into production
  • Directed the launch of three multi-tenant SaaS platforms; ran architecture reviews and established AI/ML working practices
  • Mentored an agile team of 8 engineers: seniors ship v1, juniors iterate under review
  • Led quantitative-markets decision support research: cross-asset anomaly detection and price forecasting

Senior Software Engineer

Pacific Cloud Computing Ltd.

Hong KongJanuary 2015 - November 2021

  • Built an enterprise document management SaaS with real-time collaboration for corporate clients
  • Built e-commerce platform capabilities including multi-currency payments; the SEO programme reached top search rankings
  • Pioneered the firm's blockchain product work: smart-contract products, NFT marketplace support and DeFi analytics

Senior Product Manager (Technical) / Web & Content Manager

Groupon.com

Hong KongApril 2013 - December 2014

  • Led platform modernisation from a monolithic to a service-oriented architecture as technical product manager, coordinating engineering delivery and business stakeholders
  • Introduced an A/B testing and SEO programme that delivered 150% traffic growth
  • Ran digital product operations and commercial prioritisation for the marketplace: merchandising, pricing structure and merchant coordination in an Agile environment

Senior Operations Manager

SoManyCall Telecom

Hong KongMarch 2008 - April 2013

  • Oversaw strategic and operational delivery for custom software solutions in the telecoms sector, achieving a 25% annual growth rate
  • Led a multi-disciplinary team across the full project lifecycle

Core Technical Competencies

AI-Assisted Engineering

Claude Code
Codex CLI
Gemini CLI
Cursor
Spec-Driven Delivery
Eval Harness
Three Review Gates
Agent Output Review & Debugging
Responsibility Tags

Generative AI & LLM Systems

Multi-Provider Routing
RAG
Eval Harness
Streaming SSE
Prompt Registry
Guardrails
Function Calling
Token Tiering
Self-Hosted Models

Integrations & Payments

UK Open Banking
Stripe
Firebase / Firestore
Usage-Based Billing
Credit Packs
Webhooks & Idempotency
FIX
Groq / Cerebras LPU

Cloud, DevOps & Infrastructure

Cloudflare
AWS
GCP
Docker
Pages / Functions / Workers
D1 / R2
Kubernetes
Terraform
CI/CD
OpenTelemetry
Multi-Region Failover

Backend, APIs & Microservices

TypeScript
Next.js
FastAPI
Node.js
React
REST APIs
Microservices
Event-Driven
Edge Runtime

Data Pipelines & Storage

PostgreSQL
Redis
TimescaleDB
pgvector
Firestore
Cloudflare D1 / R2
Batch Pipelines
Real-Time Pipelines

Machine Learning & Deep Learning

Python
PyTorch
Scikit-learn
XGBoost
TensorFlow
Transformers
Decision Trees
LSTM/GRU
Fine-tuning
Feature Engineering

MLOps & Production ML

Model Evaluation
A/B Testing
Model Serving
Automated Retraining
Canary Deployment
Drift Monitoring
ONNX
Triton
Quantisation

Selected Deployments

Qorinix Inference Control Plane

AI agents under my review
Problem
The board wanted "instant" AI responses for real-time agents, voice and trading alerts, inside a latency budget, a provider cost ceiling and with failover when any single provider degrades.
What shipped
A multi-tier model router across 12 providers (OpenAI, Anthropic, Gemini, DeepSeek, Qwen, Groq, Cerebras and others) with streaming SSE, a per-call cost ledger, usage-based billing and entitlements; token tiering and self-hosted models for non-latency-critical work.
Outcome and measurement
Targets sub-200ms time-to-first-token (p50) and 100-200+ tokens/sec, measured by TTFT p50/p95 per provider in the LLM Arena benchmark. Token tiering cut API spend for non-latency-critical work (no percentage published).
My role
Designed and shipped with AI agents under my review

LaSpend Open Banking

AI agents under my review
Problem
A read-only AI money assistant over FCA-regulated UK Open Banking data: subscription detection, waste scoring and assisted cancellation, with PII minimisation and a rule that it never moves money.
What shipped
Integrations with the Open Banking API, Stripe and Cloudflare Pages, Functions and D1; human approval for every user-triggered action, deterministic fallback when models fail and verified action receipts.
Outcome and measurement
Live in production. Measured by detection precision on recurring spend and cancellation completion rate (user numbers not published).
My role
Designed and shipped with AI agents under my review

Fixxmi AI Lead Matching

AI agents under my review
Problem
A Swiss consumer service marketplace needed AI lead matching across 12+ categories in DE-CH/EN within GDPR and nFADP, with Firestore europe-west6 data residency.
What shipped
Next.js 15, Firebase Cloud Functions and Firestore (europe-west6) with intent classification and matching inside the governance loop; pay-per-lead monetisation with Stripe credit packs.
Outcome and measurement
Live. Measured by match acceptance rate (volume numbers not published).
My role
Designed and shipped with AI agents under my review

Governance Loop

Built
Problem
Four products with different risk profiles needed one control layer that a finance-led board and a regulator could both read, with every decision reconstructable after the fact.
What shipped
A loop applied to every AI call: policy check, evidence-pack retrieval, prompt registry, model routing, action runtime, audit ledger, human-review escalation and policy feedback.
Outcome and measurement
Measured by audit completeness and escalation rate; controls tighten from observed production behaviour.
My role
Built

Portfolio Assistant (philip.pm)

Built
Problem
A public portfolio chatbot that answers hiring questions instantly, cannot be hijacked and costs almost nothing to run: the live, inspectable proof of hands-on work.
What shipped
Next.js edge runtime on Cloudflare Pages, a five-provider cascade with timeouts and fallback, signed-cookie rate limiting with staged daily budgets, a curated preset answer bank streamed instantly and a hardened system prompt.
Outcome and measurement
Live at www.philip.pm. Measured by fallback rate and p95 response time.
My role
Built

Education

MSc Computer Science, AI & Data Science

University of Wolverhampton, UK

2023 - 2025 | Grade: Merit

  • Dissertation: LLM-Augmented High-Frequency Trading Strategy Development
  • AI-assisted three-gate validation workflow: 85% reduction in strategy development time
  • Modules: Deep Machine Learning, Intelligent Agents, Data Science & Mining, Applying AI, Cloud Computing, Research Methods

Bachelor of Business Administration

Hong Kong University of Science & Technology

1995

  • Marketing with Information Systems minor

Selected Certifications

MLOps SpecializationGoogle Cloud | 2024
Gemini Certified EducatorGoogle | 2025-2028
CS50x Computer ScienceHarvard | 2024
SFC Licensing Exams (Papers 1/7/8/12)HKSI, Hong Kong | 2021

Open to Principal Forward Deployed, Applied AI and AI Solutions Engineering roles

Full UK right to work | London, hybrid or remote UK | Languages: English, Cantonese, Mandarin

Book a 30-min chat

Last Updated: September 2026 | Portfolio: www.Philip.pm | References available upon request