{"observation":{"id":"7466b188-0aed-468f-a6fd-245507d99fcc","tool":"askyourdatabase","tool_name":"AskYourDatabase","criterion":"plain-english-query-handling","criterion_name":"Plain English Query Handling","criterion_definition":"Can the tool understand business questions without SQL?","criterion_evidence_type":"transformation","criterion_rank_role":"decisive","criterion_rank_role_reason":"This is the core of the ranking: the tool must understand a business question without the user writing SQL. (3 of 3 judges)","scenario":"best-customers-with-unpaid-order-and-payment-method-follow-ups","scenario_name":"Best customers with unpaid-order and payment-method follow-ups","group_tag":"live-database-plain-english-queries","scenario_description":"A conversational multi-table customer analysis with follow-up questions. It asks for the best customers by both order volume and spend, then drills into unpaid orders for the top 3 and their usual payment methods. Designed to test ranking logic, join-heavy analysis, and follow-up context retention.","modality":"text","input_text":"Who are my best customers — the ones who order the most and spend the most?\n\nFollow-up 1: For the top 3 from that list — do any of them have unpaid orders?\n\nFollow-up 2: What payment methods do these top 3 usually use?","input_artifact_refs":[],"stresses":["ambiguous business term interpretation","multi-table joins","aggregation and ranking","follow-up context retention","scoping to a selected subset","payment behavior analysis"],"verdict":"worked","score":null,"score_total":null,"note":"It can interpret informal ranking language like 'the ones who order the most and spend the most' and launch the analysis directly.","evidence_state":"verified","source":null,"artifacts":[{"url":"https://cdn.futuresmart.ai/public/aidemos/f53605d222624c199b365f4beaf646f5.png?v=1","role":"output","alt":null},{"url":"https://cdn.futuresmart.ai/public/aidemos/ec0ed41fe3dd45b5b13b082d174be6b7.png?v=1","role":"output","alt":null}],"run_id":"af2abc96-3311-484b-a19d-854a2fdd2bf3","study_title":"Query Live Databases Using Plain English with AI","study_kind":"generation","research_task":"86b9y6c99","tested_at":null,"completeness":"input-and-output","input":{"state":"text","text":"Who are my best customers — the ones who order the most and spend the most?\n\nFollow-up 1: For the top 3 from that list — do any of them have unpaid orders?\n\nFollow-up 2: What payment methods do these top 3 usually use?","files":[],"modality":"text","stresses":["ambiguous business term interpretation","multi-table joins","aggregation and ranking","follow-up context retention","scoping to a selected subset","payment behavior analysis"]},"tool_page_slug":"askyourdatabase","tool_url":"https://aidemos.com/tools/askyourdatabase","permalink":"https://aidemos.com/evidence/7466b188-0aed-468f-a6fd-245507d99fcc","api_url":"https://ai.aidemos.com/v1/observations/7466b188-0aed-468f-a6fd-245507d99fcc"},"peers":[{"id":"2bd93ade-7def-4250-ad0a-dd4f54aafd4a","tool":"basedash","tool_name":"Basedash","verdict":"worked","score":null,"score_total":null,"note":"It correctly handled a conversational, multi-part customer question without requiring SQL, producing two rankings plus follow-up answers.","artifact_count":1,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/b32adc19ccfc434bb534d9c3684fee55.png?v=1","evidence_url":"https://aidemos.com/evidence/2bd93ade-7def-4250-ad0a-dd4f54aafd4a"},{"id":"736c4cc8-4b83-445e-b321-4bf1c78cd426","tool":"definite","tool_name":"Definite","verdict":"mixed","score":null,"score_total":null,"note":"It accepted the natural-language request but only ranked customers by total spend, so it did not fully understand the combined 'order the most and spend the most' intent.","artifact_count":2,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/7bb8ca311aef4c5c9a664a5f1df2814c.mp4?v=1","evidence_url":"https://aidemos.com/evidence/736c4cc8-4b83-445e-b321-4bf1c78cd426"},{"id":"32f69b4e-ce68-451c-ae6d-098434171c49","tool":"draxlr","tool_name":"Draxlr","verdict":"worked","score":null,"score_total":null,"note":"It accepted the informal best-customers request and both follow-up questions across the three-turn conversation.","artifact_count":5,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/ab41b8340d0a45188a25417725b1cede.png?v=1","evidence_url":"https://aidemos.com/evidence/32f69b4e-ce68-451c-ae6d-098434171c49"},{"id":"56eb0ed0-bf0d-4077-9f04-bd1a290baeb9","tool":"futuresmart-nl2sql-agent","tool_name":"FutureSmart NL2SQL Agent","verdict":"worked","score":null,"score_total":null,"note":"Accepts the best-customers question in plain English and returns a ranked answer without requiring SQL from the user.","artifact_count":1,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/9f6a37e2aa8044c5a801b7b0c1fe9ef9.png?v=1","evidence_url":"https://aidemos.com/evidence/56eb0ed0-bf0d-4077-9f04-bd1a290baeb9"},{"id":"78125c79-ab09-4a31-9c8e-246e8474d64b","tool":"querio","tool_name":"Querio","verdict":"worked","score":null,"score_total":null,"note":"Handles an informal conversational request and splits 'order the most' and 'spend the most' into two separate ranking dimensions instead of guessing.","artifact_count":1,"thumbnail":"https://cdn.futuresmart.ai/public/aidemos/364651fb80d04b7299109e6465d39e8d.png?v=1","evidence_url":"https://aidemos.com/evidence/78125c79-ab09-4a31-9c8e-246e8474d64b"}],"other_criteria":[{"id":"1e61c241-7f0d-4a12-977f-2e2b437b2f80","criterion":"ambiguity-handling","criterion_name":"Ambiguity Handling","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"It handles the ambiguous 'best customers' request by surfacing both order-count and spend rankings instead of silently choosing one metric.","artifact_count":1,"evidence_url":"https://aidemos.com/evidence/1e61c241-7f0d-4a12-977f-2e2b437b2f80"},{"id":"239f04d5-19bb-454b-bebc-39db63df6c63","criterion":"business-insight","criterion_name":"Business Insight","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"It adds interpretation such as Rahul Sharma being the all-rounder and Mohan Vishe being a payment-risk red flag, rather than just listing rows.","artifact_count":4,"evidence_url":"https://aidemos.com/evidence/239f04d5-19bb-454b-bebc-39db63df6c63"},{"id":"71867d8f-4b4e-4263-b1df-786dd813435f","criterion":"follow-up-context","criterion_name":"Follow-Up Context","rank_role":"context","verdict":"worked","score":null,"score_total":null,"note":"It preserves the selected top 3 across turns and reuses the earlier result set for the payment-method follow-up without issuing a new SQL query.","artifact_count":3,"evidence_url":"https://aidemos.com/evidence/71867d8f-4b4e-4263-b1df-786dd813435f"},{"id":"6a130b11-3ffe-4dd1-b405-30c887af3d65","criterion":"result-readability","criterion_name":"Result Readability","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"It presents the answer as clearly labeled ranking tables and customer-level payment summaries, with visual risk cues for unpaid or shipped orders.","artifact_count":4,"evidence_url":"https://aidemos.com/evidence/6a130b11-3ffe-4dd1-b405-30c887af3d65"},{"id":"c81647fc-befc-4195-abe0-94c81a794ac3","criterion":"sql-generation","criterion_name":"SQL Generation","rank_role":"decisive","verdict":"worked","score":null,"score_total":null,"note":"It can generate follow-up SQL constrained to exactly the top 3 customers, using hardcoded customer IDs in the WHERE clause.","artifact_count":1,"evidence_url":"https://aidemos.com/evidence/c81647fc-befc-4195-abe0-94c81a794ac3"}],"appears_in":[{"page_type":"ranking","slug":"ai-database-query-tools","title":"Best AI Tools to Query Live Databases Using Plain English","url":"https://aidemos.com/best/ai-database-query-tools","binding":"run"},{"page_type":"ranking","slug":"text-to-sql-tools","title":"Best AI Tools to Query Live Databases Using Plain English","url":"https://aidemos.com/best/text-to-sql-tools","binding":"study"},{"page_type":"use-case","slug":"query-live-databases","title":"Query Live Databases Using Plain English with AI","url":"https://aidemos.com/use-cases/query-live-databases","binding":"study"}],"same_scenario":[{"id":"2bd93ade-7def-4250-ad0a-dd4f54aafd4a","tool":"basedash","tool_name":"Basedash","verdict":"worked","score":null,"score_total":null,"note":"It correctly handled a conversational, multi-part customer question without requiring SQL, producing two rankings plus follow-up answers."},{"id":"736c4cc8-4b83-445e-b321-4bf1c78cd426","tool":"definite","tool_name":"Definite","verdict":"mixed","score":null,"score_total":null,"note":"It accepted the natural-language request but only ranked customers by total spend, so it did not fully understand the combined 'order the most and spend the most' intent."},{"id":"32f69b4e-ce68-451c-ae6d-098434171c49","tool":"draxlr","tool_name":"Draxlr","verdict":"worked","score":null,"score_total":null,"note":"It accepted the informal best-customers request and both follow-up questions across the three-turn conversation."},{"id":"56eb0ed0-bf0d-4077-9f04-bd1a290baeb9","tool":"futuresmart-nl2sql-agent","tool_name":"FutureSmart NL2SQL Agent","verdict":"worked","score":null,"score_total":null,"note":"Accepts the best-customers question in plain English and returns a ranked answer without requiring SQL from the user."},{"id":"78125c79-ab09-4a31-9c8e-246e8474d64b","tool":"querio","tool_name":"Querio","verdict":"worked","score":null,"score_total":null,"note":"Handles an informal conversational request and splits 'order the most' and 'spend the most' into two separate ranking dimensions instead of guessing."}]}