Principal Forward Deployed Engineer · Weidmann & Cie. AG · Applied AI in Financial Services

PHILIP CHEUNG

I am the engineer who gets AI from a working demo into a business that has to live with it.

Principal Forward Deployed Engineer at Weidmann & Cie. AG, a Swiss corporate finance advisory firm: a five-person team, four products in production (Qorinix, LaSpend, Fixxmi and an AI-native trading desk) and weekly reporting to a finance-led board.

MSc Computer Science, AI & Data Science (Merit)

  • Discovery to production ownership: SLAs, audit trail, rollback and on-call, not demos
  • Integration depth: Open Banking, Stripe, Firebase, Cloudflare, AWS, FIX and 12 LLM providers
  • Governance built in: policy gates, evidence packs, human-review escalation, per-call cost and audit ledgers
  • Adoption: thin working prototypes and live dashboards instead of slide decks; built with AI coding agents under a three-gate review
Philip Cheung

Professional Summary

I am the engineer who gets AI from a working demo into a business that has to live with it. As Principal Forward Deployed Engineer at Weidmann & Cie. AG, a Swiss corporate finance advisory firm, I take four products from discovery to production: Qorinix (ultra-low-latency AI inference cloud), LaSpend (read-only AI money assistant over UK Open Banking), Fixxmi (Swiss consumer service marketplace) and an AI-native trading desk. With a five-person team I own architecture, integrations and production accountability, and I report weekly to a finance-led board.

What that looks like in practice: integrating regulated Open Banking data, Stripe, Firebase and Cloudflare; routing across twelve LLM providers with a per-call cost ledger and failover; building governance into every call (policy gates, evidence packs, audit ledger, human review for high-risk actions); and winning decisions with thin working prototypes and live dashboards rather than slide decks.

I build with AI coding agents under a three-gate review and evaluation process, an approach I first validated in my MSc research (85% reduction in strategy development time). I know the systems well enough to review what the agents produce, debug them in production and be accountable for the result: I carry the pager when it breaks.

Before engineering leadership I spent years on the client side of technology: technical product management at Groupon and enterprise account management for a marketing-technology platform, working with government and multinational accounts across Hong Kong, Singapore, Taiwan and Macau. MSc Computer Science, AI and Data Science (Merit). Full UK right to work, based near London, hybrid or remote.

Core Capabilities

Discovery to Production Ownership

SLAs, audit trail, rollback and on-call; the escalation point for production incidents, not demos

Integration Depth

Open Banking, Stripe, Firebase/Firestore, Cloudflare (Pages, Functions, Workers, D1, R2), AWS, FIX, 12 LLM providers, Groq/Cerebras LPU inference

Governance Built In

Policy gates, evidence packs, human-review escalation, per-call cost ledger, audit ledger

Adoption and Stakeholder Work

Thin working prototype plus live dashboard instead of slide decks; weekly board reporting; client-side background in account management and technical product management

Deployment Leadership

Five-person team now, 8 engineers mentored at Pacific Cloud; seniors ship v1, juniors iterate under review

How I Deliver (AI-assisted)

Claude Code, Codex CLI, Gemini CLI, Cursor; eval harness and three review gates; I review, debug and own what the agents produce

Notable Achievements

Production Accountability Across Four Products

Owns end-to-end architecture and production accountability for Qorinix, LaSpend, Fixxmi and an AI-native trading desk with a five-person team, and is the escalation point for production incidents. Reports weekly to a finance-led board.

Qorinix Inference Control Plane

Designed and shipped, with AI coding agents under a three-gate review and evaluation process, a multi-tier model router across 12 providers with streaming SSE, per-call cost ledger and usage-based billing; targets sub-200ms time-to-first-token (p50).

Regulated Integrations in Production

LaSpend live over UK Open Banking, Stripe and Cloudflare D1, read-only by policy with human approval for every user-triggered action. Fixxmi AI lead matching live across 12+ categories in DE-CH/EN within GDPR and nFADP boundaries.

Governance Loop on Every AI Call

Built the loop applied to every AI call across the products: policy check, evidence-pack retrieval, prompt registry, model routing, audit ledger and human-review escalation for high-risk cases.

Hands-on, Inspectable Evidence

Built the 12-provider LLM Arena benchmark and this site's portfolio assistant by hand; both run live on philip.pm. 30+ published Lab posts on inference, agents, evaluation and delivery.

Deployment Leadership

Leads a five-person team today and mentored an agile team of 8 engineers at Pacific Cloud, where he directed the launch of three multi-tenant SaaS platforms: seniors ship v1, juniors iterate under review.

Research Foundation

MSc Computer Science, AI and Data Science (Merit). Dissertation validated the AI-assisted three-gate workflow he now uses in production, cutting strategy development time by 85%.

Commercial Track Record

150% traffic growth at Groupon through an A/B testing and SEO programme, and 25% annual growth leading operations at SoManyCall.

Professional Experience

Principal Forward Deployed Engineer

Weidmann & Cie. AG

London, UK (Hybrid)

Dec 2024 - Present

Swiss corporate finance advisory firm building an applied-AI product portfolio: Qorinix (ultra-low-latency AI inference cloud), LaSpend (read-only AI money assistant over UK Open Banking), Fixxmi (Swiss consumer service marketplace) and an AI-native trading desk. I take each product from discovery to production with a five-person team and report weekly to a finance-led board.

What I own and ship

  • Own end-to-end architecture and production accountability across all four products on AWS, Cloudflare (Pages, Functions, Workers, D1, R2) and Firebase; I am the escalation point for production incidents
  • Designed and shipped the Qorinix inference control plane with AI coding agents under a three-gate review and evaluation process: multi-tier model router across 12 providers (OpenAI, Anthropic, Gemini, DeepSeek, Qwen, Groq, Cerebras and others), streaming SSE, per-call cost ledger, usage-based billing and entitlements; targets sub-200ms time-to-first-token
  • Integrated LaSpend with regulated UK Open Banking data, Stripe and Cloudflare D1: subscription detection, waste scoring and assisted cancellation, designed read-only by policy with human approval for every user-triggered action and deterministic fallback when models fail
  • Delivered Fixxmi's AI lead matching (Next.js 15, Firebase Cloud Functions, Firestore europe-west6) across 12+ service categories in DE-CH/EN within GDPR and nFADP boundaries, with pay-per-lead Stripe credit packs
  • Built the governance loop applied to every AI call: policy check, evidence-pack retrieval, prompt registry, model routing, audit ledger and human-review escalation for high-risk cases
  • Ran a 12-provider inference benchmark (LLM Arena, live at philip.pm) to ground model, latency and cost decisions; introduced token tiering and self-hosted models for non-latency-critical workloads to cut API spend
  • Translate board, investor and regulatory requirements into scoped, dated deliveries; win decisions with thin working prototypes and live dashboards rather than slide decks

Skills, Integrations & Technologies

Forward Deployed Engineering
Open Banking
Stripe
Firebase / Firestore
Cloudflare (Pages, Workers, D1, R2)
AWS
12 LLM Providers
Groq / Cerebras
Streaming SSE
Evals & Governance
AI Coding Agents
Board Reporting
AI Engineer & Technical Architect

Pacific Cloud Computing Ltd.

Hong Kong & Remote UK

Dec 2021 - Dec 2024

Promoted internally from Senior Software Engineer to lead the firm's AI transformation across SaaS, analytics and market-intelligence products.

What I own and ship

  • Built and operated a real-time ML inference service (feature pipelines, model serving, monitoring) for the firm's analytics products, with latency and accuracy gates defined before each release
  • Implemented enterprise retrieval (RAG) and market-intelligence NLP pipelines, including the evaluation methodology used to accept them into production
  • Directed the launch of three multi-tenant SaaS platforms; ran architecture reviews and established AI/ML working practices
  • Mentored an agile team of 8 engineers: seniors ship v1, juniors iterate under review
  • Led quantitative-markets decision support research: cross-asset anomaly detection and price forecasting

Skills, Integrations & Technologies

Real-time ML Inference
Model Serving
RAG
NLP
MLOps
Multi-tenant SaaS
Architecture Review
Mentoring
Senior Software Engineer

Pacific Cloud Computing Ltd.

Hong Kong

Jan 2015 - Nov 2021

What I own and ship

  • Built an enterprise document management SaaS with real-time collaboration for corporate clients
  • Built e-commerce platform capabilities including multi-currency payments; the SEO programme reached top search rankings
  • Pioneered the firm's blockchain product work: smart-contract products, NFT marketplace support and DeFi analytics

Skills, Integrations & Technologies

Enterprise SaaS
E-commerce
Payments
Blockchain
SEO
Senior Product Manager (Technical) / Web & Content Manager

Groupon.com

Hong Kong

Apr 2013 - Dec 2014

What I own and ship

  • Led platform modernisation from a monolithic to a service-oriented architecture as technical product manager, coordinating engineering delivery and business stakeholders.
  • Ran digital product operations and commercial prioritisation for the marketplace: merchandising, pricing structure and merchant coordination in an Agile environment.

Key Outcomes

  • Introduced an A/B testing and SEO programme that delivered 150% traffic growth.

Skills, Integrations & Technologies

Technical Product Management
A/B Testing
SEO
Agile
Marketplace Operations
Senior Operations Manager

SoManyCall Telecom

Hong Kong

Mar 2008 - Apr 2013

What I own and ship

  • Oversaw strategic and operational delivery for custom software solutions in the telecoms sector.
  • Led a multi-disciplinary team across the full project lifecycle.

Key Outcomes

  • Achieved a 25% annual growth rate.

Skills, Integrations & Technologies

Operations Leadership
Telecoms
Team Leadership
Principal Forward Deployed · Case Studies

Selected Deployments

Five systems taken from discovery to production. Each case study states the problem, the constraints, what I designed, the integrations, the governance and how the outcome was measured. The role label says who did the work.

Inference Infra

Qorinix Inference Control Plane

The board wanted instant AI responses for real-time agents, voice and trading alerts within a latency budget, provider cost and failover constraints. I designed a multi-tier router across 12 providers with streaming SSE, a per-call cost ledger and usage-based billing. Measured by TTFT p50/p95 per provider in the LLM Arena benchmark: targets sub-200ms TTFT, and token tiering cut API spend for non-latency-critical work.

My role: Designed and shipped with AI agents under my review
Read the case study
Regulated Fintech

LaSpend over UK Open Banking

A read-only AI money assistant on FCA-regulated Open Banking data, with PII minimisation and a hard rule that it never moves money. Integrations: Open Banking API, Stripe, Cloudflare Pages, Functions and D1. Governance: human approval for every user-triggered action, deterministic fallback and verified action receipts. Live in production, measured by recurring-spend detection precision and cancellation completion rate.

My role: Designed and shipped with AI agents under my review
Read the case study
Consumer Marketplace

Fixxmi AI Lead Matching

A Swiss consumer marketplace needed AI lead matching across 12+ categories in DE-CH and EN, inside GDPR and nFADP boundaries with Firestore europe-west6 data residency. Built on Next.js 15, Firebase Cloud Functions and Firestore, with pay-per-lead Stripe credit packs. Live, measured by match acceptance rate.

My role: Designed and shipped with AI agents under my review
Read the case study
Governance

The Governance Loop on Every AI Call

One loop applied to every AI call across the products: policy check, evidence-pack retrieval, prompt registry, model routing, action runtime, audit ledger, human-review escalation and policy feedback. Every call is versioned, costed and replayable. Measured by audit completeness and escalation rate.

My role: Built
Read the case study
Live on this site

This Site's Portfolio Assistant

The philip.pm chatbot itself, the live and inspectable proof of hands-on work: Next.js edge runtime on Cloudflare Pages, a five-provider cascade (Groq, Cerebras, Gemini, DeepSeek, OpenRouter) with timeouts and fallback, signed-cookie rate limiting with staged daily budgets, a curated preset answer bank streamed instantly and a hardened system prompt. Measured by fallback rate and p95 response time.

My role: Built
Read the case study

PUBLIC BUILD LOG

AI Engineering Lab

Published LinkedIn posts as build-in-public evidence: inference and serving, agents and harnesses, retrieval and data, evaluation loops, and field notes from taking AI to production. Every post ships with a short video walkthrough.

30+ published postsVideo-backed

Skills & Expertise

Forward Deployed Engineering

  • Discovery, Scoping & Dated DeliveriesExpert
  • Integration Engineering in Regulated EnvironmentsExpert
  • Production Ownership (SLAs, Rollback, On-call)Advanced
  • Thin Working Prototypes & Live DashboardsExpert
  • Board, Investor & Regulator CommunicationExpert

LLM & Agent Systems

  • Multi-Provider Routing, Failover & Cost LedgerExpert
  • RAG & Evidence-Pack RetrievalAdvanced
  • Agentic Workflows & Tool UseAdvanced
  • Streaming SSE & Latency Budgets (TTFT p50/p95)Advanced
  • Token Tiering & Self-Hosted ModelsAdvanced

Governance & Evaluation

  • Eval Harness & Release GatesExpert
  • Policy Gates & Human-Review EscalationExpert
  • Audit Ledger & Replayable CallsAdvanced
  • PII Minimisation & Data Residency (GDPR, nFADP)Advanced
  • Provider & Vendor EvaluationAdvanced

Financial Services Domain

  • Open Banking & Regulated Financial DataAdvanced
  • Corporate Finance & M&A WorkflowsWorking
  • Quantitative Markets & TradingAdvanced
  • Fintech Operations & BillingAdvanced
  • Risk, Compliance & Regulated ProcessesAdvanced

Integrations

  • Open Banking APIAdvanced
  • Stripe (Subscriptions, Credit Packs, Webhooks)Advanced
  • Firebase, Firestore & Cloud FunctionsAdvanced
  • Cloudflare (Pages, Functions, Workers, D1, R2)Expert
  • AWS, FIX & Groq/Cerebras LPU InferenceWorking

How I Deliver (AI-assisted)

  • Claude Code, Codex CLI, Gemini CLI, CursorExpert
  • Spec First, Agents Generate, I Review & AcceptExpert
  • Three Review Gates & Evaluation Before MergeExpert
  • Debugging & Owning What the Agents ProduceExpert
  • Smallest Slice That Proves ValueExpert

Core Technical Competencies · AI & ML

  • ML & Deep Learning (Python, PyTorch, Scikit-learn, XGBoost, Transformers)Advanced
  • Production ML (Model Serving, Monitoring, A/B Testing, Drift)Advanced
  • Generative AI & LLM Systems (RAG, Orchestration, Guardrails, Eval Harness)Expert
  • Inference Benchmarking (TTFT, Tokens/sec, Cost per Task)Advanced

Core Technical Competencies · Platform & Data

  • Backend & APIs (TypeScript, Next.js 15, Node.js, FastAPI)Advanced
  • Cloud & Infrastructure (AWS, GCP/Firebase, Cloudflare, Docker, Kubernetes)Advanced
  • Data & Storage (PostgreSQL, Redis, TimescaleDB, pgvector, Firestore, D1)Working
  • Streaming / SSE, Edge Runtime & Event-Driven PatternsAdvanced

Education & Certifications

Education

MSc in Computer Science (AI & Data Science, Merit)
2023 - 2025

AI Systems, Applied ML, and LLM Decision Workflows

University of Wolverhampton

Visit Website

Focused on production-oriented AI, machine learning system design, and evidence-based decision support architecture.

Completed
Bachelor of Business Administration (BBA)
1995

Marketing (with IS minor)

The Hong Kong University of Science & Technology

Visit Website

Foundation in business strategy, communication, and cross-functional leadership.

Completed

Certifications & Qualifications

Gemini Certified Educator
2025

Applied AI enablement and practical adoption coaching.

Google Cloud MLOps Specialization
2024

Google Cloud

Visit Website

Production lifecycle practices for ML systems.

Harvard CS50x
2024

Harvard University

Visit Website

Computer science foundations and software engineering discipline.

Licensing Examination for Securities and Futures Intermediaries (LE)
2021

SFC / HKSI

Visit Website

Passed Papers 1, 7, 8, 12.

Frequently Asked Questions

The questions hiring managers ask about a Principal Forward Deployed Engineer

Resources & Playbooks

Forward Deployment 90-Day Playbook
Playbook2026
Author: Philip Cheung
philip.pm

How I take an AI system from first customer conversation to a handed-over production service in 90 days: discovery, scoping, integration, governance, adoption and handover, with the gates and artefacts for each stage.

AI Leadership 90-Day Execution Framework
Playbook2026
Author: Philip Cheung
philip.pm

A practical framework for first-90-day AI leadership: baseline assessment, platform foundation, and high-value use-case delivery plan.

LLM Platform Governance Design
Playbook2026
Author: Philip Cheung
philip.pm

Policy-first architecture for LLM deployments, covering prompt controls, human-review escalation, and auditability requirements.

Evaluation and Release Gates for AI Systems
Playbook2026
Author: Philip Cheung
philip.pm

How to set measurable release criteria, benchmark drift, and maintain quality over time in production AI systems.

Schedule an Interview
Book a time slot that works for you using my Calendly scheduling system.

Loading booking calendar...

If the embedded widget is blank, please use direct booking.

Open Calendly

Contact Me

Email

Use contact form only

Phone

Use contact form only

Location

London, United Kingdom