We evaluate generative AI applications, assistants, RAG systems, and AI-enabled product features against defined business and quality criteria. Engagements may include evaluation strategy, golden datasets, human review, LLM-as-a-judge, factuality and faithfulness checks, relevance, safety, consistency, retrieval quality, tool-use validation, and regression testing. Frameworks such as Ragas, DeepEval, and LangSmith can be used where appropriate. The outcome is a repeatable evaluation process that helps teams compare changes, detect regressions, and make better release decisions for nondeterministic systems.