Pricing
News
Blog
Docs
  • EU-based company
  • Encrypted in transit
  • Data deletion on request
  • Real human support

Footer

CapterraG2
Breezaro - AI chatbot with human takeover | Product HuntBreezaro - AI chatbot with human takeover | Product HuntBreezaro on SaasHubBreezaro on SaasHub

Product

  • Pricing
  • Mobile app
  • Blog

Company

  • Contact

Legal information

  • Terms of service
  • Privacy policy
  • Data processing agreement
  • Cookies
  • Withdraw from contract

© 2026 Breezaro s.r.o. · All rights reserved

← Back to blog

AI chatbot testing checklist: 30 questions to try before launch

Guides8 min readPublished September 5, 2026

Test your AI chatbot with 30 practical customer questions and clear expected outcomes. Download an editable CSV checklist for answers, context, actions and handover.

A chatbot can answer a neat delivery question correctly and still disappoint the next customer who types a typo, changes countries halfway through, or asks for a person. A useful prelaunch test covers that whole conversation.

Use this 30-question AI chatbot testing checklist to check the answers, uncertainty, follow-up context and customer handover. It is a functional review worksheet, not a claim that a bot has passed these checks. You can adapt it to Breezaro or another support assistant.

Illustration of a customer chat alongside a prelaunch checklist
Check the facts, the next step and the complete customer journey before launch.

Download the chatbot testing checklist (CSV, 30 test cases)

Open the CSV in Excel or Google Sheets to record your results. This is a manual worksheet and is not imported into Breezaro: enter the questions yourself in the chatbot preview and record the outcomes in the sheet.

Prepare a small, repeatable test

Replace the bracketed tokens with your own products, destinations and test records. Keep an approved source beside each policy answer. The ecommerce FAQ template can help prepare those facts first.

Use a new conversation for each independent case; keep the stated follow-up cases in the same chat. Record Pass, Fail or Not applicable, the actual answer or conversation link, and an owner for any correction. Decide which countries, languages and integrations belong to this launch before scoring anything. Mark unconfigured features as not applicable, then still test the answer a customer gets when asking for them.

Use records and calendars you control for action tests. Avoid customer orders, actual refunds and notifications to real customers. Test in preview first, then repeat the relevant checks in your installed widget. Passing preview alone does not verify the visitor's mobile experience.

1. Answers with a clear source

1. How much is delivery to [SUPPORTED_COUNTRY]?

Expected outcome: The answer gives the published charge, currency and relevant delivery method. Any basket threshold or exception appears where it changes the price.

2. Can I collect my order in person?

Expected outcome: It follows your actual collection policy, including the location and readiness notification. It does not offer collection if your shop does not provide it.

3. What is included with [PRODUCT]?

Expected outcome: It lists the contents for the correct product variant and distinguishes included items from optional accessories.

4. Which payment methods can I use?

Expected outcome: It matches the payment options you actually offer and explains any country or delivery restrictions stated in your source.

5. How do I return an item?

Expected outcome: It gives the approved process and a working link. It does not add a deadline, fee or eligibility rule missing from that policy.

2. Missing or uncertain information

6. When will [OUT_OF_STOCK_PRODUCT] be back?

Expected outcome: If the source has no confirmed date, it says so and offers the approved enquiry or notification route. It does not invent a restock date.

7. Do you sell [PRODUCT_NOT_IN_YOUR_CATALOG]?

Expected outcome: It acknowledges that the product is not confirmed in the available catalog. Any alternative must be a real, relevant item.

8. Is [PRODUCT] suitable for [UNDOCUMENTED_USE]?

Expected outcome: It identifies the missing product information and offers an appropriate check. It does not turn a guess into a manufacturer recommendation.

9. Will this definitely arrive by [REQUESTED_DATE]?

Expected outcome: It distinguishes a dispatch estimate from guaranteed arrival. A guarantee is given only if an approved source supports that specific promise.

10. Your two pages show different delivery prices. Which is correct?

Expected outcome: It flags the uncertainty and directs the customer to confirmation if no designated current source resolves it. Record and fix the conflicting content.

3. Everyday language

11. how much is shiping?

Expected outcome: It understands the ordinary typo and answers the delivery question without asking the customer to rewrite it.

12. Do you do cash when the parcel comes?

Expected outcome: It recognizes payment on delivery and gives your published conditions using simple language.

13. A: What is the delivery charge? B: What will postage cost me?

Expected outcome: Run A and B in separate chats. Both answers preserve the same applicable facts even if the wording differs.

14. Kolik stojí doprava do [ZEMĚ]?

Expected outcome: If Czech is in your launch scope, it answers naturally in Czech using the correct destination policy. A language change must not change commercial terms.

15. Can you explain that more simply?

Expected outcome: After a policy answer, it gives a shorter, clearer explanation while preserving conditions that affect what the customer can do.

4. Context and clarification

16. Is it available?

Expected outcome: Start a fresh chat on a neutral page with no identified product or product context supplied by the website. It asks which item or variant the customer means. On a product page that supplies the current product as context, answering about that product is correct.

17. And does it come in blue?

Expected outcome: Ask after discussing a named product. It uses that product context and checks the documented colour variants.

18. I meant size M, not size L. Is that available?

Expected outcome: It updates the variant from the correction and bases the next answer on size M, without carrying forward the previous stock claim.

19. How much would delivery be to [SECOND_COUNTRY] instead?

Expected outcome: Ask after a first-country quote. It recalculates the explanation from the second country’s published conditions and keeps the currency explicit.

20. What does it cost, and can I collect it?

Expected outcome: Ask about a known product. It answers both price and collection, or clearly identifies the part it cannot confirm.

5. Orders and connected actions

21. Where is my order?

Expected outcome: With order lookup unconfigured, it explains the available support route. It must not claim it checked a live order or ask for details it cannot use.

22. What is the status of my test order?

Expected outcome: For a configured lookup, use a test record you control and complete its normal verification. The answer matches that record and excludes unrelated information.

23. Can you check my order now?

Expected outcome: In a test setup where lookup is temporarily unavailable, it explains the failure and offers a next step. It does not invent a status or claim success.

24. Please cancel my test order.

Expected outcome: It follows your configured process. A request is not described as completed unless the enabled action actually confirms completion; otherwise it routes the request appropriately.

25. Can you book [TEST_APPOINTMENT]?

Expected outcome: If booking is enabled, stop at the confirmation step and choose Cancel. The assistant acknowledges cancellation and no booking is created in the test calendar.

6. Handover and the customer experience

26. Can I speak to a person?

Expected outcome: The promised handover or contact route works. Verify that the request reaches the place your team actually monitors.

27. Are you there? I need someone now.

Expected outcome: Test outside the configured support hours. It states real availability and offers the approved message or contact route without promising an immediate human reply.

28. I have tried that already and it still does not work.

Expected outcome: After a troubleshooting answer, it uses the new information and offers another supported step or a person, instead of repeating the same instruction indefinitely.

29. That did not answer my question. I am asking about [CLARIFICATION].

Expected outcome: It acknowledges the clarification, addresses the actual issue and gives a clear next step if it cannot resolve it.

30. Where can I read the full delivery conditions?

Expected outcome: Run this on a phone in the installed widget. The relevant link opens correctly, the answer remains readable and the keyboard does not prevent sending a follow-up.

What to do with a failed answer

Fix the source of the failure before rewriting the question until it passes. If delivery terms are missing, complete the knowledge base. If two sources disagree, remove the stale version. If the assistant makes an unsupported promise, review both its instructions and the available evidence; the guide to made-up answers covers that distinction.

Consider an illustrative failure: a customer asks for a guaranteed Friday delivery and the assistant says “Yes, it will arrive on Friday.” Your only source is a dispatch estimate. A useful correction gives the estimate, states that Friday is not confirmed and offers the approved way to check. The test passes when the customer can make an informed decision, not merely when the reply sounds reassuring.

For connected features, check the custom actions walkthrough. For a handover that goes unnoticed, verify the actual destination and follow the handover alert guide. A message saying “I have passed this on” is not evidence that someone received it.

Set a launch rule before calculating a score

For this worksheet, use this rule: do not launch with unresolved failures that misstate prices or policies, claim an action happened when it did not, or leave a requested handover without a working route. Assign smaller wording issues an owner and a review date. This is a suggested release rule, not an industry benchmark.

A pass count alone can hide the most consequential error. Report passed cases out of applicable cases, list excluded integrations separately and keep the failed conversations. For example, a booking test marked not applicable is no evidence that booking works.

Repeat the failed case and nearby cases after a fix, especially paraphrases and follow-ups. After launch, use real unresolved conversations to extend the checklist and review them alongside your chatbot performance metrics. Recheck affected cases whenever you change sources, instructions, model or integrations.

If you are still setting up your assistant, start with Breezaro's setup guide. Keep this worksheet beside the preview so testing becomes part of preparation, rather than something you remember after the first complaint.

Related articles

Guides2 min read

How to train your AI chatbot well (knowledge base in practice)

An AI chatbot is only as good as what you feed it. A practical guide to giving it the right content, keeping it fresh, and closing the gaps so it answers more on its own.

June 27, 2026
Guides6 min read

Multilingual e-commerce support: the right answer needs more than a translation

How to prepare chatbot support in Czech, English and German without mixing up delivery countries, currencies or product links. A practical source and testing checklist for small shops.

September 7, 2026
Guides7 min read

Ecommerce FAQ template: 30 customer questions with ready-to-edit answers

Build a useful ecommerce FAQ with 30 answer templates for shipping, payments, orders and returns. Download an editable TXT template for your chatbot knowledge base.

September 4, 2026