The Modern AI Data Stack
The AI data stack: data warehouse (Snowflake, BigQuery) → transformation (dbt) → feature store → model training/fine-tuning → vector database → LLM API. You don't need to build all of this — but you should understand it.
- Snowflake Cortex embeds LLM capabilities directly in the data warehouse.
- dbt (data build tool) is the standard for transforming raw data into AI-ready features.
- Vector databases store embeddings — the mathematical representation of text meaning.
---
Data Quality for AI
AI amplifies data quality problems. Biased training data produces biased models. Incomplete data produces unreliable predictions. The 'garbage in, garbage out' principle is more important than ever.
- Poor data quality costs UK businesses £25.6B annually (IBCS 2024).
- For RAG: 80% of ingestion effort is cleaning and chunking source documents.
- Data labelling quality: a single mislabelled example can corrupt a fine-tuned model.
---
Privacy-Preserving AI
Federated learning trains AI without centralising sensitive data. Differential privacy adds noise to protect individual records. These techniques let you build AI on sensitive data legally.
- Federated learning: Apple uses it for keyboard prediction — no data leaves the device.
- Synthetic data generation (Gretel, Mostly AI) creates GDPR-compliant training data.
- On-device AI (Apple Intelligence) processes sensitive queries without sending to the cloud.
---
Cloud AI Platforms
AWS Bedrock, Azure OpenAI Service, and Google Vertex AI provide enterprise-grade AI infrastructure: security, compliance, SLAs, and cost management. If you're on one of the big three clouds, start here.
- AWS Bedrock provides access to Claude, Llama, Mistral, and others in one API.
- Azure OpenAI Service is HIPAA-compliant and SOC 2 certified.
- Google Vertex AI integrates with BigQuery for data-to-AI pipelines without moving data.
---
On-Premises vs Cloud AI
On-prem AI (running models on your own hardware) is necessary for classified data, financial regulations, or data-sovereign requirements. The cost has fallen: an RTX 4090 runs a 70B model for £1,500.
- Llama 3 70B runs at 30 tokens/second on a single A100 GPU.
- On-prem vector DB (pgvector in Postgres) handles most enterprise RAG needs.
- Hybrid: cloud for development and burst capacity, on-prem for production sensitive workloads.
---
The Context Window: Opportunity and Cost
The context window is how much text an AI can process in one call. Claude 3.5 handles 200K tokens (~500 pages). This enables document-scale analysis but at 10–100× the cost of shorter prompts.
- 200K context = processing a 500-page report in a single API call.
- Long-context calls cost $3–15 per million tokens (model-dependent).
- Cache frequently used context (system prompt, shared docs) to reduce cost 90%.
---
AI Architecture Patterns
The four patterns: single LLM call (fastest, cheapest), chain (sequential LLM steps), parallel (multiple LLMs simultaneously), and agentic loop (LLM + tools + memory). Match pattern to problem.
- Chain: research → summarise → draft → review — best for multi-step document work.
- Parallel: run 3 different AI models on the same query and vote on the answer.
- Agentic loop: the AI decides what to do next until the task is complete.
---
Edge AI and Mobile AI
AI is moving to the device: Apple Intelligence, Google Gemini Nano, Qualcomm AI chips. This enables AI without internet, with zero latency, and without data leaving the device — a major privacy and performance advantage.
- Apple Intelligence processes requests on-device for privacy-sensitive tasks.
- Gemini Nano 1 runs on a Pixel 9 with 2GB RAM.
- Edge AI enables real-time inference: language translation, object detection, anomaly monitoring.
Sign in to track your progress and earn a certificate.
Sign in