AI & LLM Systems Pod
Collaborative cohort following Andrej Karpathy and Fast.ai roadmaps. We build fine-tuned small LLMs and RAG pipelines.
Take turns: One student acts as interviewer while the other answers using the STAR technique (Situation, Task, Action, Result).
Explain the architectural difference between self-attention and cross-attention in Transformer models.
Discuss queries, keys, and values projection matrices. Detail where encoder-decoder models inject cross-attention.
How do you mitigate catastrophic forgetting when fine-tuning an open-source model using LoRA / QLoRA?
Explain rank matrices (A & B), parameter efficiency, freezing base weights, and validation loss tracking.
Describe a project where you encountered vanishing or exploding gradients and how you diagnosed it.
Use the STAR framework: Situation (training loss stagnated), Task (debug tensor values), Action (gradient clipping, layer norm), Result.
How do you evaluate retrieval precision and context hallucination in a production RAG system?
Mention context recall, context precision, answer relevancy, and embedding vector similarity thresholds.
What excites you most about open-weights AI models versus proprietary closed APIs?
Highlight latency control, local data privacy, cost predictability, and independence from cloud vendor lock-in.