Learn LLM Path / pillar 9 of 10
Sign in to track this pillar and save notes. The resource list stays open to everyone.
stop guess-and-tweak. Evaluation-driven development is the single biggest predictor of agent-building success (per Andrew Ng)
Objective vs LLM-judge; error analysis; traces KEY
Agentic AI - Module 4 (Andrew Ng)articlein course
RAG evals (RAGAS): faithfulness, precision, recall
RAGAS in code (run in CI)
RAGAS Evaluation Tutorial (local, no API) - write-uparticleread
Tracing/observability in practice
LangSmith docsdocsdocs