Quick summary Evaluate actions and answers: polished text can hide a failed workflow. Check three surfaces: the outcome, tool-use path and final external state. Combine suitable graders: deterministic checks, rubrics and human review. Learn from failures: turn traces into regression cases and measure cost and constraints. AI agent evaluation starts with a simple reality: an […]
Interview question:Uber predicts that ride demand in a particular area will increase significantly in the next 20 minutes. How would you design an Uber demand prediction system that forecasts demand and proactively encourages drivers to move to that area? This is a classic marketplace and machine-learning system design question. The interviewer is testing whether you […]