Testing AI prompts and comparing models with promptfoo
How to automate AI prompt testing and compare models using promptfoo for reliable outputs.
How to automate AI prompt testing and compare models using promptfoo for reliable outputs.
Article on testing AI chatbots, including prompt injection, permissions, and bypassing guardrails using influence techniques.
Using GitHub Agentic Workflows with AI to test Bicep deployments for best practices, region emissions, and latency.
How to avoid the AI expertise trap by testing AI in areas you know well to accurately assess its capabilities.
Testing OpenAI's ChatGPT Images 2.0 with a Where's Waldo-style raccoon and ham radio prompt, comparing results across models.
Explores three common but misleading answers QA professionals give about AI's impact on testing, using Maslow's Hammer as a critique.
Explains AI evals: automated checks for non-deterministic AI outputs using LLMs to score against expectations, not exact matches.
Explores the unique challenges of testing Generative AI and Large Language Models, contrasting it with traditional software testing approaches.
Introduces AlignEval, an app for building and automating LLM evaluators, making the process easier and more data-driven.
Call for participation in WAIT #3, a peer conference on AI in software testing, seeking experienced testers to share and evaluate real-world AI testing experiences.