Principal Forward Deployed Engineer · Weidmann & Cie. AG · Applied AI in Financial Services
PHILIP CHEUNG
I am the engineer who gets AI from a working demo into a business that has to live with it.
Principal Forward Deployed Engineer at Weidmann & Cie. AG, a Swiss corporate finance advisory firm: a five-person team, four products in production (Qorinix, LaSpend, Fixxmi and an AI-native trading desk) and weekly reporting to a finance-led board.
MSc Computer Science, AI & Data Science (Merit)
- Discovery to production ownership: SLAs, audit trail, rollback and on-call, not demos
- Integration depth: Open Banking, Stripe, Firebase, Cloudflare, AWS, FIX and 12 LLM providers
- Governance built in: policy gates, evidence packs, human-review escalation, per-call cost and audit ledgers
- Adoption: thin working prototypes and live dashboards instead of slide decks; built with AI coding agents under a three-gate review
Professional Summary
I am the engineer who gets AI from a working demo into a business that has to live with it. As Principal Forward Deployed Engineer at Weidmann & Cie. AG, a Swiss corporate finance advisory firm, I take four products from discovery to production: Qorinix (ultra-low-latency AI inference cloud), LaSpend (read-only AI money assistant over UK Open Banking), Fixxmi (Swiss consumer service marketplace) and an AI-native trading desk. With a five-person team I own architecture, integrations and production accountability, and I report weekly to a finance-led board.
What that looks like in practice: integrating regulated Open Banking data, Stripe, Firebase and Cloudflare; routing across twelve LLM providers with a per-call cost ledger and failover; building governance into every call (policy gates, evidence packs, audit ledger, human review for high-risk actions); and winning decisions with thin working prototypes and live dashboards rather than slide decks.
I build with AI coding agents under a three-gate review and evaluation process, an approach I first validated in my MSc research (85% reduction in strategy development time). I know the systems well enough to review what the agents produce, debug them in production and be accountable for the result: I carry the pager when it breaks.
Before engineering leadership I spent years on the client side of technology: technical product management at Groupon and enterprise account management for a marketing-technology platform, working with government and multinational accounts across Hong Kong, Singapore, Taiwan and Macau. MSc Computer Science, AI and Data Science (Merit). Full UK right to work, based near London, hybrid or remote.
Core Capabilities
Discovery to Production Ownership
SLAs, audit trail, rollback and on-call; the escalation point for production incidents, not demos
Integration Depth
Open Banking, Stripe, Firebase/Firestore, Cloudflare (Pages, Functions, Workers, D1, R2), AWS, FIX, 12 LLM providers, Groq/Cerebras LPU inference
Governance Built In
Policy gates, evidence packs, human-review escalation, per-call cost ledger, audit ledger
Adoption and Stakeholder Work
Thin working prototype plus live dashboard instead of slide decks; weekly board reporting; client-side background in account management and technical product management
Deployment Leadership
Five-person team now, 8 engineers mentored at Pacific Cloud; seniors ship v1, juniors iterate under review
How I Deliver (AI-assisted)
Claude Code, Codex CLI, Gemini CLI, Cursor; eval harness and three review gates; I review, debug and own what the agents produce
Notable Achievements
Production Accountability Across Four Products
Owns end-to-end architecture and production accountability for Qorinix, LaSpend, Fixxmi and an AI-native trading desk with a five-person team, and is the escalation point for production incidents. Reports weekly to a finance-led board.
Qorinix Inference Control Plane
Designed and shipped, with AI coding agents under a three-gate review and evaluation process, a multi-tier model router across 12 providers with streaming SSE, per-call cost ledger and usage-based billing; targets sub-200ms time-to-first-token (p50).
Regulated Integrations in Production
LaSpend live over UK Open Banking, Stripe and Cloudflare D1, read-only by policy with human approval for every user-triggered action. Fixxmi AI lead matching live across 12+ categories in DE-CH/EN within GDPR and nFADP boundaries.
Governance Loop on Every AI Call
Built the loop applied to every AI call across the products: policy check, evidence-pack retrieval, prompt registry, model routing, audit ledger and human-review escalation for high-risk cases.
Hands-on, Inspectable Evidence
Built the 12-provider LLM Arena benchmark and this site's portfolio assistant by hand; both run live on philip.pm. 30+ published Lab posts on inference, agents, evaluation and delivery.
Deployment Leadership
Leads a five-person team today and mentored an agile team of 8 engineers at Pacific Cloud, where he directed the launch of three multi-tenant SaaS platforms: seniors ship v1, juniors iterate under review.
Research Foundation
MSc Computer Science, AI and Data Science (Merit). Dissertation validated the AI-assisted three-gate workflow he now uses in production, cutting strategy development time by 85%.
Commercial Track Record
150% traffic growth at Groupon through an A/B testing and SEO programme, and 25% annual growth leading operations at SoManyCall.
Professional Experience
Weidmann & Cie. AG
London, UK (Hybrid)
Swiss corporate finance advisory firm building an applied-AI product portfolio: Qorinix (ultra-low-latency AI inference cloud), LaSpend (read-only AI money assistant over UK Open Banking), Fixxmi (Swiss consumer service marketplace) and an AI-native trading desk. I take each product from discovery to production with a five-person team and report weekly to a finance-led board.
What I own and ship
- Own end-to-end architecture and production accountability across all four products on AWS, Cloudflare (Pages, Functions, Workers, D1, R2) and Firebase; I am the escalation point for production incidents
- Designed and shipped the Qorinix inference control plane with AI coding agents under a three-gate review and evaluation process: multi-tier model router across 12 providers (OpenAI, Anthropic, Gemini, DeepSeek, Qwen, Groq, Cerebras and others), streaming SSE, per-call cost ledger, usage-based billing and entitlements; targets sub-200ms time-to-first-token
- Integrated LaSpend with regulated UK Open Banking data, Stripe and Cloudflare D1: subscription detection, waste scoring and assisted cancellation, designed read-only by policy with human approval for every user-triggered action and deterministic fallback when models fail
- Delivered Fixxmi's AI lead matching (Next.js 15, Firebase Cloud Functions, Firestore europe-west6) across 12+ service categories in DE-CH/EN within GDPR and nFADP boundaries, with pay-per-lead Stripe credit packs
- Built the governance loop applied to every AI call: policy check, evidence-pack retrieval, prompt registry, model routing, audit ledger and human-review escalation for high-risk cases
- Ran a 12-provider inference benchmark (LLM Arena, live at philip.pm) to ground model, latency and cost decisions; introduced token tiering and self-hosted models for non-latency-critical workloads to cut API spend
- Translate board, investor and regulatory requirements into scoped, dated deliveries; win decisions with thin working prototypes and live dashboards rather than slide decks
Skills, Integrations & Technologies
Pacific Cloud Computing Ltd.
Hong Kong & Remote UK
Promoted internally from Senior Software Engineer to lead the firm's AI transformation across SaaS, analytics and market-intelligence products.
What I own and ship
- Built and operated a real-time ML inference service (feature pipelines, model serving, monitoring) for the firm's analytics products, with latency and accuracy gates defined before each release
- Implemented enterprise retrieval (RAG) and market-intelligence NLP pipelines, including the evaluation methodology used to accept them into production
- Directed the launch of three multi-tenant SaaS platforms; ran architecture reviews and established AI/ML working practices
- Mentored an agile team of 8 engineers: seniors ship v1, juniors iterate under review
- Led quantitative-markets decision support research: cross-asset anomaly detection and price forecasting
Skills, Integrations & Technologies
Pacific Cloud Computing Ltd.
Hong Kong
What I own and ship
- Built an enterprise document management SaaS with real-time collaboration for corporate clients
- Built e-commerce platform capabilities including multi-currency payments; the SEO programme reached top search rankings
- Pioneered the firm's blockchain product work: smart-contract products, NFT marketplace support and DeFi analytics
Skills, Integrations & Technologies
Groupon.com
Hong Kong
What I own and ship
- Led platform modernisation from a monolithic to a service-oriented architecture as technical product manager, coordinating engineering delivery and business stakeholders.
- Ran digital product operations and commercial prioritisation for the marketplace: merchandising, pricing structure and merchant coordination in an Agile environment.
Key Outcomes
- Introduced an A/B testing and SEO programme that delivered 150% traffic growth.
Skills, Integrations & Technologies
SoManyCall Telecom
Hong Kong
What I own and ship
- Oversaw strategic and operational delivery for custom software solutions in the telecoms sector.
- Led a multi-disciplinary team across the full project lifecycle.
Key Outcomes
- Achieved a 25% annual growth rate.
Skills, Integrations & Technologies
Selected Deployments
Five systems taken from discovery to production. Each case study states the problem, the constraints, what I designed, the integrations, the governance and how the outcome was measured. The role label says who did the work.
Qorinix Inference Control Plane
The board wanted instant AI responses for real-time agents, voice and trading alerts within a latency budget, provider cost and failover constraints. I designed a multi-tier router across 12 providers with streaming SSE, a per-call cost ledger and usage-based billing. Measured by TTFT p50/p95 per provider in the LLM Arena benchmark: targets sub-200ms TTFT, and token tiering cut API spend for non-latency-critical work.
LaSpend over UK Open Banking
A read-only AI money assistant on FCA-regulated Open Banking data, with PII minimisation and a hard rule that it never moves money. Integrations: Open Banking API, Stripe, Cloudflare Pages, Functions and D1. Governance: human approval for every user-triggered action, deterministic fallback and verified action receipts. Live in production, measured by recurring-spend detection precision and cancellation completion rate.
Fixxmi AI Lead Matching
A Swiss consumer marketplace needed AI lead matching across 12+ categories in DE-CH and EN, inside GDPR and nFADP boundaries with Firestore europe-west6 data residency. Built on Next.js 15, Firebase Cloud Functions and Firestore, with pay-per-lead Stripe credit packs. Live, measured by match acceptance rate.
The Governance Loop on Every AI Call
One loop applied to every AI call across the products: policy check, evidence-pack retrieval, prompt registry, model routing, action runtime, audit ledger, human-review escalation and policy feedback. Every call is versioned, costed and replayable. Measured by audit completeness and escalation rate.
This Site's Portfolio Assistant
The philip.pm chatbot itself, the live and inspectable proof of hands-on work: Next.js edge runtime on Cloudflare Pages, a five-provider cascade (Groq, Cerebras, Gemini, DeepSeek, OpenRouter) with timeouts and fallback, signed-cookie rate limiting with staged daily budgets, a curated preset answer bank streamed instantly and a hardened system prompt. Measured by fallback rate and p95 response time.
PUBLIC BUILD LOG
AI Engineering Lab
Published LinkedIn posts as build-in-public evidence: inference and serving, agents and harnesses, retrieval and data, evaluation loops, and field notes from taking AI to production. Every post ships with a short video walkthrough.
Skills & Expertise
Forward Deployed Engineering
- Discovery, Scoping & Dated DeliveriesExpert
- Integration Engineering in Regulated EnvironmentsExpert
- Production Ownership (SLAs, Rollback, On-call)Advanced
- Thin Working Prototypes & Live DashboardsExpert
- Board, Investor & Regulator CommunicationExpert
LLM & Agent Systems
- Multi-Provider Routing, Failover & Cost LedgerExpert
- RAG & Evidence-Pack RetrievalAdvanced
- Agentic Workflows & Tool UseAdvanced
- Streaming SSE & Latency Budgets (TTFT p50/p95)Advanced
- Token Tiering & Self-Hosted ModelsAdvanced
Governance & Evaluation
- Eval Harness & Release GatesExpert
- Policy Gates & Human-Review EscalationExpert
- Audit Ledger & Replayable CallsAdvanced
- PII Minimisation & Data Residency (GDPR, nFADP)Advanced
- Provider & Vendor EvaluationAdvanced
Financial Services Domain
- Open Banking & Regulated Financial DataAdvanced
- Corporate Finance & M&A WorkflowsWorking
- Quantitative Markets & TradingAdvanced
- Fintech Operations & BillingAdvanced
- Risk, Compliance & Regulated ProcessesAdvanced
Integrations
- Open Banking APIAdvanced
- Stripe (Subscriptions, Credit Packs, Webhooks)Advanced
- Firebase, Firestore & Cloud FunctionsAdvanced
- Cloudflare (Pages, Functions, Workers, D1, R2)Expert
- AWS, FIX & Groq/Cerebras LPU InferenceWorking
How I Deliver (AI-assisted)
- Claude Code, Codex CLI, Gemini CLI, CursorExpert
- Spec First, Agents Generate, I Review & AcceptExpert
- Three Review Gates & Evaluation Before MergeExpert
- Debugging & Owning What the Agents ProduceExpert
- Smallest Slice That Proves ValueExpert
Core Technical Competencies · AI & ML
- ML & Deep Learning (Python, PyTorch, Scikit-learn, XGBoost, Transformers)Advanced
- Production ML (Model Serving, Monitoring, A/B Testing, Drift)Advanced
- Generative AI & LLM Systems (RAG, Orchestration, Guardrails, Eval Harness)Expert
- Inference Benchmarking (TTFT, Tokens/sec, Cost per Task)Advanced
Core Technical Competencies · Platform & Data
- Backend & APIs (TypeScript, Next.js 15, Node.js, FastAPI)Advanced
- Cloud & Infrastructure (AWS, GCP/Firebase, Cloudflare, Docker, Kubernetes)Advanced
- Data & Storage (PostgreSQL, Redis, TimescaleDB, pgvector, Firestore, D1)Working
- Streaming / SSE, Edge Runtime & Event-Driven PatternsAdvanced
Education & Certifications
Education
AI Systems, Applied ML, and LLM Decision Workflows
University of Wolverhampton
Visit WebsiteFocused on production-oriented AI, machine learning system design, and evidence-based decision support architecture.
Marketing (with IS minor)
The Hong Kong University of Science & Technology
Visit WebsiteFoundation in business strategy, communication, and cross-functional leadership.
Certifications & Qualifications
Applied AI enablement and practical adoption coaching.
Google Cloud
Visit WebsiteProduction lifecycle practices for ML systems.
Harvard University
Visit WebsiteComputer science foundations and software engineering discipline.
SFC / HKSI
Visit WebsitePassed Papers 1, 7, 8, 12.
Frequently Asked Questions
The questions hiring managers ask about a Principal Forward Deployed Engineer
Resources & Playbooks
How I take an AI system from first customer conversation to a handed-over production service in 90 days: discovery, scoping, integration, governance, adoption and handover, with the gates and artefacts for each stage.
A practical framework for first-90-day AI leadership: baseline assessment, platform foundation, and high-value use-case delivery plan.
Policy-first architecture for LLM deployments, covering prompt controls, human-review escalation, and auditability requirements.
How to set measurable release criteria, benchmark drift, and maintain quality over time in production AI systems.
Loading booking calendar...
If the embedded widget is blank, please use direct booking.
Open CalendlyContact Me
Use contact form only
Phone
Use contact form only
Location
London, United Kingdom
