← Leaderboard
Google: Gemini 3.5 Flash
google/gemini-3.5-flash · google · context 1 048 576 · in $1.50/1M · out $9.00/1M
Global Index
802
95% CI [756–849] · index v0.2.0
Per-domain scores
| Domain | Score (95% CI) | Accuracy (IRT) | Consistency | Calibration | Contam. Δ | p50 | $/1k | |
|---|---|---|---|---|---|---|---|---|
| agentic | 641 [535–746] | 0.695 | 0.86 | 0.89 | 0.192 | 1.8s | $28.75 | |
| code | 880 [763–996] | 0.800 | 1.00 | 1.00 | 0.000 | 1.7s | $17.20 | |
| instruction following | 845 [724–967] | 0.793 | 0.80 | 0.99 | 0.000 | 1.6s | $14.73 | |
| knowledge | 725 [553–897] | 0.546 | 0.98 | 1.00 | 0.000 | 1.7s | $5.90 | |
| math | 846 [695–997] | 0.743 | 1.00 | 1.00 | 0.000 | 1.7s | $9.31 | |
| multilingual | 824 [662–986] | 0.706 | 1.00 | 1.00 | 0.000 | 1.6s | $5.45 | |
| reasoning | 860 [721–999] | 0.767 | 1.00 | 1.00 | 0.000 | 1.7s | $6.21 | |
| terminal | 863 [774–953] | 0.826 | 1.00 | 0.97 | 0.038 | 1.8s | $14.52 | |
| vision ocr | 735 [564–906] | 0.559 | 1.00 | 1.00 | 0.000 | 2.6s | $5.81 |
Every answer, every test
Full transparency: the model's latest graded answer on every instantiated item of the current index version. Answer keys never leave the server; anchor-item prompts are withheld to protect the longitudinal subset.
agentic 25/30 correct
correctagentic.tools.context-load-v1conf 100% · 1.5s · $0.078 · 7869 tok
question
You are an order-operations agent working strictly through tool calls.
TOOL CATALOG:
- restock(item: string, qty: int) — reorders stock for a large pending order
- cancel_order(order_id: int) — cancels a small pending order
- ship_order(order_id: int) — ships a paid order (out of scope here)
- refund(order_id: int, amount: int) — refunds a customer (out of scope here)
- notify_customer(customer: string, message: string) — sends a notification (not required by this policy)
ORDER LEDGER (247 records, format: id|customer|region|item|qty|status):
```
1540|acme|north|pump|54|pending
1872|gale|east|frame|60|held
1991|gale|east|gasket|70|shipped
1711|ionic|east|rotor|55|pending
1785|ionic|east|rotor|25|pending
1482|juno|west|frame|20|paid
1852|birch|north|sensor|24|pending
1546|acme|north|pump|49|paid
2156|juno|west|pump|69|pending
1465|ember|north|valve|47|shipped
1832|acme|north|frame|36|held
2043|dorian|east|pump|78|pending
1727|fulton|west|cable|52|held
1802|dorian|south|valve|76|pending
1623|ionic|east|pump|20|shipped
2256|ember|north|cable|67|pending
1317|ember|south|cable|17|shipped
2119|birch|east|gasket|53|paid
1531|ember|north|valve|64|pending
1401|birch|north|rotor|15|pending
2170|birch|east|rotor|45|pending
1594|birch|west|frame|60|shipped
1292|ember|south|gasket|18|pending
1452|gale|east|cable|67|paid
1986|ember|west|rotor|12|pending
1511|dorian|north|pump|35|held
2107|fulton|west|frame|84|held
1477|harbor|east|rotor|17|shipped
1820|birch|north|cable|41|pending
1558|cobalt|west|sensor|83|shipped
2165|juno|east|pump|27|shipped
1606|ember|east|panel|88|paid
1838|dorian|east|panel|89|pending
1297|ember|east|pump|22|pending
1837|dorian|south|cable|63|held
1786|ember|west|pump|34|shipped
1913|cobalt|south|sensor|95|held
2183|acme|east|panel|52|shipped
1457|ember|north|pump|20|held
1321|ember|north|valve|64|pending
1551|acme|south|pump|77|pending
1566|fulton|east|valve|45|pending
1645|harbor|west|panel|72|shipped
1463|ionic|east|cable|63|pending
1930|ember|south|frame|41|pending
1346|birch|west|frame|24|shipped
1509|juno|west|frame|14|held
2185|fulton|north|gasket|50|pending
1573|birch|east|cable|48|shipped
1585|juno|north|rotor|97|pending
2078|gale|west|sensor|52|held
2152|fulton|west|sensor|59|pending
1612|cobalt|east|gasket|36|paid
2146|gale|east|frame|77|pending
1301|ember|south|frame|20|shipped
1705|cobalt|west|rotor|88|held
1357|dorian|west|pump|82|shipped
1823|juno|north|cable|73|shipped
1431|cobalt|east|rotor|19|held
1620|birch|east|pump|42|pending
1816|fulton|west|valve|67|shipped
2239|cobalt|west|pump|92|pending
2254|cobalt|north|panel|65|shipped
1484|juno|west|pump|52|held
1271|ember|north|frame|13|pending
1633|fulton|east|panel|55|paid
1672|fulton|north|rotor|19|held
1834|cobalt|north|rotor|80|shipped
1954|cobalt|north|frame|95|pending
2055|gale|west|frame|83|paid
1602|gale|south|sensor|28|held
1397|dorian|east|panel|97|paid
1682|dorian|east|valve|67|shipped
1281|ember|west|panel|17|pending
2097|acme|south|panel|75|paid
1765|juno|west|panel|10|pending
1827|ionic|north|pump|54|pending
1600|dorian|north|gasket|80|shipped
1889|ember|east|frame|93|shipped
2075|acme|north|rotor|63|held
1562|harbor|east|cable|82|shipped
1358|fulton|west|gasket|93|shipped
1947|fulton|west|gasket|89|held
1774|cobalt|east|rotor|73|paid
1636|juno|west|frame|25|paid
1926|cobalt|south|rotor|83|paid
1527|gale|east|pump|40|pending
1387|harbor|north|cable|41|paid
1894|gale|north|panel|46|pending
2036|dorian|west|rotor|83|paid
1470|fulton|north|cable|94|shipped
1276|ember|south|valve|87|paid
1859|birch|north|gasket|78|paid
1385|harbor|north|pump|92|shipped
2195|acme|west|pump|61|paid
1796|juno|east|sensor|54|paid
1650|acme|east|panel|94|paid
2188|harbor|west|sensor|78|held
1879|cobalt|south|panel|50|paid
2227|ionic|south|rotor|61|pending
2134|acme|south|gasket|74|pending
1864|fulton|south|frame|28|held
2200|acme|north|rotor|29|shipped
1339|juno|west|valve|41|pending
1943|acme|east|sensor|35|held
1937|ember|east|cable|85|held
2247|juno|east|sensor|96|held
1671|ember|west|gasket|22|paid
1903|ionic|south|frame|45|paid
1393|cobalt|west|frame|32|shipped
1898|birch|south|frame|30|paid
1628|fulton|west|rotor|68|shipped
2245|fulton|south|frame|78|held
1303|ember|south|rotor|80|pending
2176|fulton|north|valve|83|held
2234|birch|west|cable|38|paid
2112|cobalt|west|pump|43|paid
2096|acme|west|pump|39|held
1858|ionic|west|rotor|72|paid
1319|ember|south|valve|56|pending
1793|ionic|north|rotor|68|pending
2092|ionic|north|sensor|79|held
1348|fulton|north|sensor|86|held
2069|birch|east|cable|21|paid
1679|gale|north|panel|15|paid
2225|juno|north|pump|95|paid
2045|ember|north|gasket|20|shipped
2064|harbor|east|frame|83|held
2139|fulton|south|pump|15|held
2009|dorian|west|cable|66|pending
1403|juno|east|cable|53|pending
2051|juno|north|rotor|40|pending
1328|gale|south|sensor|75|held
1410|harbor|east|sensor|31|held
1590|dorian|west|valve|95|shipped
1354|ionic|west|sensor|80|held
1528|harbor|east|cable|73|pending
1699|juno|north|cable|96|shipped
2084|ember|west|pump|57|pending
1424|fulton|west|gasket|70|paid
2218|gale|east|rotor|37|pending
1582|harbor|south|pump|99|paid
1265|ember|south|cable|37|pending
1660|ember|north|pump|41|pending
1285|ember|south|cable|75|shipped
1732|birch|west|frame|63|pending
1905|juno|south|gasket|99|paid
1658|birch|east|panel|92|pending
1378|birch|south|cable|22|held
1461|harbor|west|frame|86|pending
1595|ionic|south|sensor|48|paid
1437|birch|south|frame|70|pending
1310|ember|west|frame|14|pending
1738|fulton|east|pump|69|paid
1578|ionic|east|frame|19|held
1667|ember|east|pump|63|shipped
1767|acme|south|valve|60|shipped
1981|ionic|west|gasket|14|shipped
2196|dorian|east|pump|83|paid
1692|acme|west|valve|74|held
1536|acme|east|cable|16|paid
1490|dorian|east|frame|14|shipped
1277|ember|south|pump|39|pending
1927|juno|north|rotor|44|shipped
1797|birch|east|gasket|54|held
1998|ionic|west|valve|28|held
1883|dorian|west|frame|34|shipped
1412|gale|west|sensor|28|held
2012|harbor|west|valve|76|held
2074|acme|north|panel|81|pending
2019|cobalt|south|gasket|67|shipped
2167|birch|north|rotor|41|shipped
1688|ionic|west|cable|29|shipped
1805|acme|east|pump|36|shipped
1514|ember|east|panel|55|shipped
1537|birch|south|pump|87|held
1749|juno|east|frame|15|shipped
2091|gale|south|pump|65|held
1740|dorian|west|frame|26|shipped
1745|ionic|west|rotor|51|shipped
1865|acme|north|rotor|25|paid
1944|gale|south|panel|19|held
2217|juno|south|cable|41|shipped
1809|acme|north|frame|90|paid
1974|harbor|south|valve|66|pending
1702|ionic|west|rotor|32|held
1835|juno|west|rotor|67|shipped
2101|ionic|west|valve|30|held
1953|birch|north|frame|68|shipped
1494|juno|north|valve|64|shipped
1691|birch|south|gasket|23|shipped
2131|gale|west|panel|54|shipped
2109|gale|east|sensor|55|pending
1657|cobalt|south|rotor|61|shipped
2127|acme|east|gasket|22|paid
1729|birch|north|sensor|31|shipped
1718|dorian|north|frame|79|shipped
1840|birch|east|frame|52|shipped
1372|birch|west|panel|94|pending
1964|ember|east|valve|68|held
1491|dorian|north|valve|69|held
1758|dorian|east|pump|25|shipped
1366|dorian|east|cable|71|paid
1364|birch|east|pump|88|shipped
1617|gale|north|gasket|57|held
2205|gale|west|pump|19|shipped
1756|birch|east|valve|15|held
1781|cobalt|south|frame|44|held
1323|ember|south|cable|38|paid
1722|dorian|east|gasket|64|held
2032|dorian|west|cable|63|pending
2021|cobalt|west|valve|82|paid
1957|harbor|north|cable|69|paid
1912|gale|west|cable|99|paid
2088|birch|east|gasket|36|paid
1501|cobalt|south|rotor|35|held
1436|juno|north|rotor|67|pending
2141|birch|east|rotor|16|shipped
2160|cobalt|east|frame|85|held
1530|harbor|south|gasket|92|pending
1938|juno|north|panel|64|shipped
1335|acme|north|cable|68|paid
1446|fulton|west|sensor|47|shipped
2061|acme|north|rotor|41|held
2180|fulton|south|gasket|58|shipped
2211|juno|east|panel|18|held
2157|cobalt|west|cable|57|shipped
1784|ionic|north|panel|78|held
1920|juno|west|pump|53|shipped
1753|gale|north|frame|39|held
1338|dorian|west|valve|73|held
1521|birch|east|panel|78|shipped
1343|ember|east|gasket|85|held
1969|ionic|east|panel|17|held
2028|birch|west|rotor|30|shipped
1502|acme|north|frame|71|shipped
1417|cobalt|south|gasket|22|paid
1804|birch|north|frame|29|paid
1440|ionic|north|frame|75|shipped
1780|dorian|north|cable|24|held
2014|ionic|north|pump|23|shipped
2123|ionic|west|frame|50|pending
2004|ember|west|pump|36|paid
2003|cobalt|south|valve|72|pending
1640|birch|east|rotor|89|paid
1935|birch|west|pump|52|shipped
1847|ionic|east|valve|65|held
```
POLICY (apply exactly):
- Consider ONLY orders where customer = "ember" AND region = "south" AND status = "pending".
- Process those orders in ASCENDING order id.
- For each: if qty > 41, call restock(item, qty) with that order's item and qty; otherwise call cancel_order(order_id).
- Call nothing else. Every other record is irrelevant no matter how similar it looks.
Output the full ordered sequence of tool calls as a JSON array of {"tool": string, "args": object}.
Give your final answer inside ONE fenced code block (```), containing exactly the required content and nothing else. After the code block, add one plain-text line:
CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>model answer:
[
{"tool": "cancel_order", "args": {"order_id": 1265}},
{"tool": "cancel_order", "args": {"order_id": 1277}},
{"tool": "cancel_order", "args": {"order_id": 1292}},
{"tool": "restock", "args": {"item": "rotor", "qty": 80}},
{"tool": "restock", "args": {"item": "valve", "qty": 56}},
{"tool": "cancel_order", "args": {"order_id": 1930}}
]correctagentic.tools.ledger-v1conf 100% · 2.0s · $0.024 · 2540 tok
question
You are a treasury agent operating strictly through tool calls.
TOOL CATALOG:
- transfer(from: string, to: string, amount: int) — moves funds between accounts; FAILS if it would overdraw the source
- top_up_from_reserve(account: string, amount: int) — adds funds to an account from the corporate reserve
- get_balance(account: string) — reads a balance (you already have all balances below — reads are unnecessary and forbidden)
- freeze_account(account: string) — compliance freeze (not part of this task)
- convert_currency(account: string, currency: string) — FX conversion (all amounts are already in USD)
OPENING BALANCES:
- tango: $652
- kilo: $564
- lima: $415
PAYMENT INSTRUCTIONS (execute in exactly this order):
1. pay $335 from "lima" to "kilo"
2. pay $208 from "tango" to "lima"
3. pay $334 from "kilo" to "lima"
4. pay $298 from "tango" to "kilo"
5. pay $121 from "tango" to "lima"
RULES:
- transfer() fails on overdraft. If a payment would overdraw its source account at the moment of execution, first call top_up_from_reserve() on the source with EXACTLY the shortfall (no more, no less), then execute the transfer.
- Track balances as they change: earlier payments affect later ones.
- Never call tools that are not needed.
Output the full ordered sequence of tool calls needed to accomplish the goal, as a JSON array of objects {"tool": string, "args": object}. Use exactly the tool and argument names from the catalog. Do not call any tool that is not required.
Give your final answer inside ONE fenced code block (```), containing exactly the required content and nothing else. After the code block, add one plain-text line:
CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>model answer:
[
{
"tool": "transfer",
"args": {
"from": "lima",
"to": "kilo",
"amount": 335
}
},
{
"tool": "transfer",
"args": {
"from": "tango",
"to": "lima",
"amount": 208
}
},
{
"tool": "transfer",
"args": {
"from": "kilo",
"to": "lima",
"amount": 334
}
},
{
"tool": "transfer",
"args": {
"from": "tango",
"to": "kilo",
"amount": 298
}
},
{
"tool": "transfer",
"args": {
"from": "tango",
"to": "lima",
"amount": 121
}
}
]wrongagentic.tools.triage-v1conf 100% · 1.5s · $0.024 · 2566 tok
question
You operate a support desk strictly through tool calls.
TOOL CATALOG:
- create_ticket(title: string, priority: int) — opens a ticket; IDs are assigned sequentially: the 1st created ticket is "TCK-1", the 2nd "TCK-2", etc.
- assign(ticket_id: string, agent: string) — assigns an open ticket to an agent
- escalate(ticket_id: string) — marks a ticket as escalated
- close_ticket(ticket_id: string, resolution: string) — closes a ticket with a resolution note
- send_email(to: string, body: string) — sends an email (not part of the triage policy)
- archive_ticket(ticket_id: string) — archives a closed ticket (nightly job does this automatically)
- set_reminder(ticket_id: string, hours: int) — sets a follow-up reminder
ROUTING POLICY (apply exactly, in this order, for each incident, processing incidents in the order listed):
1. Create a ticket for the incident (title = the incident text verbatim, priority as given).
2. If the incident is an exact duplicate of an earlier incident in this list, close its ticket immediately with resolution "duplicate of <ID of the earlier ticket>" and do nothing else for it.
3. Otherwise, if priority ≥ 7, escalate the ticket BEFORE assigning it.
4. Assign the ticket to the agent responsible for its category.
CATEGORY → AGENT:
- auth → rivera
- infra → chen
- data → dubois
INCIDENTS:
1. "SSO loop on login" (category: auth, priority 7)
2. "SSO loop on login" (category: auth, priority 7)
3. "dashboard shows stale numbers" (category: data, priority 8)
4. "SSO loop on login" (category: auth, priority 6)
Output the full ordered sequence of tool calls needed to accomplish the goal, as a JSON array of objects {"tool": string, "args": object}. Use exactly the tool and argument names from the catalog. Do not call any tool that is not required.
Give your final answer inside ONE fenced code block (```), containing exactly the required content and nothing else. After the code block, add one plain-text line:
CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>model answer:
[
{
"tool": "create_ticket",
"args": {
"title": "SSO loop on login",
"priority": 7
}
},
{
"tool": "escalate",
"args": {
"ticket_id": "TCK-1"
}
},
{
"tool": "assign",
"args": {
"ticket_id": "TCK-1",
"agent": "rivera"
}
},
{
"tool": "create_ticket",
"args": {
"title": "SSO loop on login",
"priority": 7
}
},
{
"tool": "close_ticket",
"args": {
"ticket_id": "TCK-2",
"resolution": "duplicate of TCK-1"
}
},
{
"tool": "create_ticket",
"args": {
"title": "dashboard shows stale numbers",
"priority": 8
}
},
{
"tool": "escalate",
"args": {
"ticket_id": "TCK-3"
}
},
{
"tool": "assign",
"args": {
"ticket_id": "TCK-3",
"agent": "dubois"
}
},
{
"tool": "create_ticket",
"args": {
"title": "SSO loop on login",
"priority": 6
}
},
{
"tool": "close_ticket",
"args": {
"ticket_id": "TCK-4",
"resolution": "duplicate of TCK-1"
}
}
]correctagentic.tools.context-load-v1conf 100% · 1.8s · $0.092 · 9314 tok
question
You are an order-operations agent working strictly through tool calls.
TOOL CATALOG:
- restock(item: string, qty: int) — reorders stock for a large pending order
- cancel_order(order_id: int) — cancels a small pending order
- ship_order(order_id: int) — ships a paid order (out of scope here)
- refund(order_id: int, amount: int) — refunds a customer (out of scope here)
- notify_customer(customer: string, message: string) — sends a notification (not required by this policy)
ORDER LEDGER (283 records, format: id|customer|region|item|qty|status):
```
2255|ember|west|valve|84|held
1765|harbor|west|cable|27|paid
1666|cobalt|south|valve|63|paid
1568|ionic|south|rotor|11|paid
2370|dorian|east|rotor|49|paid
2140|harbor|west|gasket|61|held
1992|ionic|west|panel|51|held
1327|birch|west|panel|16|pending
1854|ember|south|panel|90|shipped
1555|dorian|east|valve|66|pending
1328|birch|east|valve|47|shipped
2165|juno|west|cable|51|held
2182|dorian|east|rotor|30|shipped
1396|ember|north|valve|47|paid
2011|cobalt|west|pump|77|pending
1826|birch|west|pump|51|paid
2166|ember|west|pump|35|held
2217|juno|west|pump|24|shipped
1403|gale|east|sensor|54|paid
2365|acme|west|frame|48|held
1841|ionic|east|cable|98|pending
1712|juno|west|valve|43|pending
1936|cobalt|south|sensor|76|pending
1464|ember|north|sensor|54|held
2187|gale|west|sensor|91|paid
1699|dorian|east|pump|21|paid
2021|fulton|west|pump|82|paid
1938|dorian|south|valve|23|held
1468|birch|east|panel|58|pending
2195|acme|south|sensor|60|shipped
1453|ember|south|pump|69|shipped
1494|ionic|west|sensor|86|paid
2371|cobalt|north|rotor|32|held
2362|ionic|west|valve|16|shipped
2283|gale|west|panel|79|pending
2388|ember|east|rotor|98|shipped
2266|birch|north|panel|21|pending
2229|juno|north|rotor|79|shipped
1656|dorian|south|sensor|53|shipped
1455|harbor|west|gasket|18|pending
1345|birch|east|cable|50|pending
2143|ember|south|rotor|58|pending
2299|gale|south|cable|91|shipped
1621|harbor|west|pump|25|paid
1813|cobalt|south|valve|66|held
2301|acme|west|panel|64|paid
2244|harbor|south|gasket|46|pending
1595|fulton|west|frame|76|paid
1907|gale|east|valve|92|pending
1692|juno|east|pump|33|held
1611|harbor|west|gasket|53|pending
2189|harbor|east|valve|57|shipped
1809|gale|north|panel|42|paid
1635|acme|south|rotor|12|held
2170|dorian|east|gasket|99|paid
1706|birch|east|panel|47|held
1601|fulton|east|pump|38|held
2063|ionic|east|gasket|84|held
1519|dorian|north|pump|48|held
1499|dorian|east|cable|90|pending
1516|acme|south|sensor|60|held
1994|birch|west|frame|14|shipped
1434|dorian|west|valve|87|paid
1782|acme|south|gasket|72|held
2038|cobalt|south|valve|19|paid
1562|cobalt|south|panel|18|paid
2288|ember|east|sensor|97|held
1717|harbor|north|frame|98|held
2111|juno|west|valve|42|pending
1974|acme|west|rotor|55|shipped
2030|harbor|east|panel|50|paid
1778|dorian|west|panel|66|shipped
2060|ionic|north|rotor|84|pending
1767|harbor|west|gasket|35|shipped
2128|ionic|south|panel|52|paid
1339|birch|south|panel|54|pending
1959|juno|north|gasket|40|held
1981|fulton|north|cable|22|paid
1822|harbor|east|cable|54|held
2220|ionic|west|rotor|70|held
2319|ember|west|gasket|86|shipped
1386|dorian|north|sensor|17|paid
2325|fulton|north|panel|25|paid
2292|dorian|east|gasket|55|held
2012|fulton|south|panel|15|paid
2279|fulton|east|frame|26|pending
1584|dorian|north|panel|16|pending
2110|cobalt|south|cable|91|held
1439|ember|east|valve|16|paid
1551|birch|north|frame|68|held
1322|birch|east|panel|28|paid
1688|fulton|east|sensor|29|held
2104|ember|north|pump|94|shipped
2331|fulton|north|valve|29|paid
2067|harbor|south|gasket|15|shipped
2275|gale|east|sensor|67|held
1354|harbor|east|sensor|66|held
1422|ionic|east|pump|53|held
2028|harbor|east|cable|58|pending
2251|cobalt|south|frame|49|shipped
2180|gale|west|gasket|98|paid
1301|birch|east|panel|70|pending
2053|gale|north|pump|96|shipped
2298|gale|south|rotor|50|pending
1444|birch|south|panel|40|paid
2112|harbor|west|gasket|56|paid
1885|harbor|south|valve|52|pending
2181|juno|east|frame|56|shipped
1318|birch|south|rotor|45|pending
1668|fulton|south|pump|39|pending
2173|gale|east|panel|98|paid
1713|dorian|south|rotor|73|pending
1662|cobalt|south|gasket|36|paid
1963|juno|north|rotor|75|paid
1927|ember|south|pump|52|shipped
1530|harbor|north|valve|81|shipped
1902|ionic|west|frame|12|paid
1867|dorian|east|cable|97|held
1952|fulton|west|rotor|57|pending
1408|cobalt|north|pump|62|shipped
1424|acme|north|cable|19|shipped
2335|harbor|east|sensor|44|pending
1610|acme|south|rotor|43|pending
1482|juno|west|pump|99|held
1689|ionic|south|frame|54|paid
2084|ionic|north|gasket|22|held
2116|ember|north|valve|31|pending
2230|acme|east|gasket|45|held
2205|dorian|north|rotor|19|held
1820|dorian|east|sensor|95|held
1931|cobalt|east|sensor|30|shipped
1802|cobalt|east|cable|67|paid
2346|ionic|north|frame|64|held
1889|harbor|south|pump|93|pending
1675|harbor|east|rotor|14|paid
1534|gale|north|sensor|22|paid
1306|birch|north|sensor|93|pending
1759|birch|east|sensor|85|held
1649|ember|east|sensor|23|held
1716|acme|south|pump|66|paid
1861|dorian|west|frame|23|held
1838|ionic|north|gasket|99|pending
1783|ionic|south|sensor|96|shipped
1877|acme|south|gasket|52|pending
2313|ionic|south|gasket|18|paid
1928|birch|west|pump|45|held
1825|cobalt|south|pump|11|held
2377|cobalt|south|frame|97|shipped
2341|juno|south|valve|48|paid
2356|ember|north|panel|59|paid
1506|dorian|south|gasket|47|paid
1565|acme|west|sensor|32|shipped
2211|gale|south|pump|85|held
2332|fulton|west|rotor|11|shipped
1915|juno|south|rotor|41|held
1641|fulton|east|sensor|99|held
2227|ionic|east|rotor|87|held
1485|birch|west|frame|38|held
1771|harbor|north|valve|53|pending
1833|gale|north|frame|33|shipped
1335|birch|east|frame|88|pending
1749|ember|north|frame|72|shipped
2035|dorian|west|gasket|22|shipped
2135|ember|east|gasket|61|paid
1734|cobalt|west|cable|77|paid
2243|dorian|north|gasket|89|shipped
1368|ionic|west|frame|91|shipped
1701|dorian|west|frame|72|paid
2056|dorian|south|panel|74|shipped
2278|gale|east|cable|68|shipped
1382|birch|west|cable|91|paid
2096|harbor|south|cable|96|held
1673|acme|east|cable|42|paid
1945|ember|north|sensor|30|pending
2122|fulton|north|valve|96|paid
1539|dorian|west|pump|50|shipped
1490|ember|east|frame|57|held
1571|fulton|south|valve|77|pending
1727|birch|west|panel|91|held
1879|acme|north|cable|48|held
2349|cobalt|south|cable|99|pending
1681|acme|east|cable|66|pending
2133|ember|west|pump|19|held
2039|gale|south|cable|81|held
2113|harbor|north|gasket|52|paid
2202|juno|south|panel|70|shipped
1566|harbor|east|rotor|22|held
2384|juno|east|rotor|83|held
2088|harbor|north|cable|96|paid
1309|birch|east|pump|91|held
1785|birch|north|pump|24|shipped
1742|birch|east|frame|30|shipped
2364|fulton|east|valve|16|pending
2121|harbor|north|frame|44|shipped
2050|gale|north|rotor|52|shipped
1996|ember|south|pump|59|pending
1852|juno|south|rotor|97|shipped
2273|acme|north|rotor|29|paid
1735|cobalt|west|valve|50|held
1787|dorian|north|valve|61|paid
1352|birch|east|sensor|47|paid
1553|cobalt|south|frame|29|paid
1977|ember|north|panel|95|shipped
2261|cobalt|north|gasket|29|shipped
1377|cobalt|east|gasket|81|held
1472|cobalt|south|sensor|89|pending
1875|ionic|south|rotor|89|shipped
1442|ember|east|panel|59|paid
2070|birch|south|rotor|51|pending
1459|harbor|west|pump|54|paid
1390|juno|west|frame|49|shipped
1429|ember|west|panel|76|paid
1633|ember|north|rotor|49|shipped
1512|fulton|south|panel|78|paid
1451|ionic|north|gasket|37|shipped
1588|gale|south|cable|55|paid
1920|birch|west|valve|50|held
1729|birch|east|rotor|37|pending
2306|ember|south|rotor|99|shipped
1545|dorian|south|gasket|84|shipped
2155|gale|north|pump|98|paid
1600|harbor|east|panel|95|shipped
1913|birch|south|valve|86|shipped
1373|cobalt|west|pump|39|pending
2149|harbor|west|valve|23|held
2291|acme|west|valve|72|paid
1360|cobalt|south|gasket|56|shipped
1380|dorian|west|panel|38|held
1988|juno|east|sensor|57|pending
1670|harbor|south|pump|60|paid
1614|acme|south|gasket|92|paid
1856|acme|east|frame|99|paid
1347|birch|north|frame|62|pending
1863|ember|east|rotor|43|shipped
2369|harbor|west|panel|41|held
1847|harbor|east|pump|53|held
1665|acme|east|sensor|50|shipped
1738|gale|east|pump|55|pending
2046|acme|south|pump|18|paid
1388|ember|west|rotor|85|paid
2098|fulton|north|cable|33|held
1794|juno|east|rotor|51|paid
2077|gale|east|cable|53|paid
1895|gale|south|pump|99|shipped
1423|ionic|south|pump|91|shipped
1344|birch|east|cable|56|held
1628|harbor|west|sensor|99|held
2239|ionic|east|sensor|44|pending
2125|gale|east|sensor|72|pending
1797|juno|west|valve|91|held
1605|ionic|south|sensor|96|pending
2160|gale|west|cable|36|held
1524|juno|north|frame|60|shipped
1823|ember|south|sensor|11|held
2094|fulton|north|gasket|73|held
2007|dorian|north|pump|43|shipped
1722|fulton|west|frame|38|pending
1690|fulton|west|gasket|97|shipped
1756|gale|west|sensor|87|paid
1314|birch|east|panel|48|pending
1707|dorian|south|valve|24|paid
1873|juno|north|cable|73|held
1829|juno|north|sensor|52|paid
1844|ember|east|sensor|81|held
2003|ember|south|cable|99|shipped
1918|dorian|north|cable|82|held
1732|juno|east|cable|85|held
1364|gale|east|valve|67|paid
1420|acme|east|gasket|12|pending
1378|gale|north|pump|86|pending
2232|birch|east|gasket|10|paid
1476|acme|south|rotor|22|held
1968|ember|west|valve|41|pending
2139|fulton|north|valve|37|held
1577|ember|west|frame|39|paid
1326|birch|east|rotor|31|pending
1647|ember|east|frame|82|paid
2100|ember|east|rotor|39|paid
2008|juno|west|panel|92|pending
2017|harbor|north|sensor|10|pending
1810|dorian|west|sensor|77|held
1415|cobalt|south|rotor|24|shipped
1855|juno|west|panel|97|paid
```
POLICY (apply exactly):
- Consider ONLY orders where customer = "birch" AND region = "east" AND status = "pending".
- Process those orders in ASCENDING order id.
- For each: if qty > 63, call restock(item, qty) with that order's item and qty; otherwise call cancel_order(order_id).
- Call nothing else. Every other record is irrelevant no matter how similar it looks.
Output the full ordered sequence of tool calls as a JSON array of {"tool": string, "args": object}.
Give your final answer inside ONE fenced code block (```), containing exactly the required content and nothing else. After the code block, add one plain-text line:
CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>model answer:
[
{"tool": "restock", "args": {"item": "panel", "qty": 70}},
{"tool": "cancel_order", "args": {"order_id": 1314}},
{"tool": "cancel_order", "args": {"order_id": 1326}},
{"tool": "restock", "args": {"item": "frame", "qty": 88}},
{"tool": "cancel_order", "args": {"order_id": 1345}},
{"tool": "cancel_order", "args": {"order_id": 1468}},
{"tool": "cancel_order", "args": {"order_id": 1729}}
]correctagentic.tools.deploy-v1conf 100% · 1.5s · $0.013 · 1345 tok
question
You are a release agent operating strictly through tool calls.
TOOL CATALOG:
- deploy(service: string) — deploys a service; FAILS if any dependency is not yet deployed
- health_check(service: string) — runs a post-deploy health probe
- rollback(service: string) — reverts a bad deploy (nothing fails in this scenario)
- scale(service: string, replicas: int) — changes replica count (out of scope)
- restart(service: string) — restarts a service (out of scope)
SERVICES AND THEIR DEPENDENCIES (a service can only be deployed after ALL its dependencies):
- gateway: notifier
- notifier: (none)
- search: notifier
- reports: notifier
POLICY:
- Deploy in waves: in each wave, deploy every service whose dependencies are already deployed, in alphabetical order; repeat until all services are deployed.
- The service "search" is flagged unstable: call health_check on it immediately after deploying it.
- Call nothing else.
Output the full ordered sequence of tool calls needed to accomplish the goal, as a JSON array of objects {"tool": string, "args": object}. Use exactly the tool and argument names from the catalog. Do not call any tool that is not required.
Give your final answer inside ONE fenced code block (```), containing exactly the required content and nothing else. After the code block, add one plain-text line:
CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>model answer:
[
{
"tool": "deploy",
"args": {
"service": "notifier"
}
},
{
"tool": "deploy",
"args": {
"service": "gateway"
}
},
{
"tool": "deploy",
"args": {
"service": "reports"
}
},
{
"tool": "deploy",
"args": {
"service": "search"
}
},
{
"tool": "health_check",
"args": {
"service": "search"
}
}
]correctagentic.tools.ledger-v1conf 100% · 1.4s · $0.025 · 2646 tok
question
You are a treasury agent operating strictly through tool calls.
TOOL CATALOG:
- transfer(from: string, to: string, amount: int) — moves funds between accounts; FAILS if it would overdraw the source
- top_up_from_reserve(account: string, amount: int) — adds funds to an account from the corporate reserve
- get_balance(account: string) — reads a balance (you already have all balances below — reads are unnecessary and forbidden)
- freeze_account(account: string) — compliance freeze (not part of this task)
- convert_currency(account: string, currency: string) — FX conversion (all amounts are already in USD)
OPENING BALANCES:
- tango: $570
- alpha: $166
- oscar: $523
PAYMENT INSTRUCTIONS (execute in exactly this order):
1. pay $97 from "tango" to "oscar"
2. pay $427 from "tango" to "oscar"
3. pay $372 from "alpha" to "tango"
4. pay $594 from "alpha" to "tango"
5. pay $592 from "oscar" to "tango"
RULES:
- transfer() fails on overdraft. If a payment would overdraw its source account at the moment of execution, first call top_up_from_reserve() on the source with EXACTLY the shortfall (no more, no less), then execute the transfer.
- Track balances as they change: earlier payments affect later ones.
- Never call tools that are not needed.
Output the full ordered sequence of tool calls needed to accomplish the goal, as a JSON array of objects {"tool": string, "args": object}. Use exactly the tool and argument names from the catalog. Do not call any tool that is not required.
Give your final answer inside ONE fenced code block (```), containing exactly the required content and nothing else. After the code block, add one plain-text line:
CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>model answer:
[
{
"tool": "transfer",
"args": {
"from": "tango",
"to": "oscar",
"amount": 97
}
},
{
"tool": "transfer",
"args": {
"from": "tango",
"to": "oscar",
"amount": 427
}
},
{
"tool": "top_up_from_reserve",
"args": {
"account": "alpha",
"amount": 206
}
},
{
"tool": "transfer",
"args": {
"from": "alpha",
"to": "tango",
"amount": 372
}
},
{
"tool": "top_up_from_reserve",
"args": {
"account": "alpha",
"amount": 594
}
},
{
"tool": "transfer",
"args": {
"from": "alpha",
"to": "tango",
"amount": 594
}
},
{
"tool": "transfer",
"args": {
"from": "oscar",
"to": "tango",
"amount": 592
}
}
]wrongagentic.tools.triage-v1conf 100% · 1.6s · $0.022 · 2401 tok
question
You operate a support desk strictly through tool calls.
TOOL CATALOG:
- create_ticket(title: string, priority: int) — opens a ticket; IDs are assigned sequentially: the 1st created ticket is "TCK-1", the 2nd "TCK-2", etc.
- assign(ticket_id: string, agent: string) — assigns an open ticket to an agent
- escalate(ticket_id: string) — marks a ticket as escalated
- close_ticket(ticket_id: string, resolution: string) — closes a ticket with a resolution note
- send_email(to: string, body: string) — sends an email (not part of the triage policy)
- archive_ticket(ticket_id: string) — archives a closed ticket (nightly job does this automatically)
- set_reminder(ticket_id: string, hours: int) — sets a follow-up reminder
ROUTING POLICY (apply exactly, in this order, for each incident, processing incidents in the order listed):
1. Create a ticket for the incident (title = the incident text verbatim, priority as given).
2. If the incident is an exact duplicate of an earlier incident in this list, close its ticket immediately with resolution "duplicate of <ID of the earlier ticket>" and do nothing else for it.
3. Otherwise, if priority ≥ 9, escalate the ticket BEFORE assigning it.
4. Assign the ticket to the agent responsible for its category.
CATEGORY → AGENT:
- data → novak
- payments → chen
- infra → tanaka
INCIDENTS:
1. "dashboard shows stale numbers" (category: data, priority 9)
2. "invoice total wrong" (category: payments, priority 5)
3. "invoice total wrong" (category: payments, priority 5)
4. "dashboard shows stale numbers" (category: data, priority 9)
Output the full ordered sequence of tool calls needed to accomplish the goal, as a JSON array of objects {"tool": string, "args": object}. Use exactly the tool and argument names from the catalog. Do not call any tool that is not required.
Give your final answer inside ONE fenced code block (```), containing exactly the required content and nothing else. After the code block, add one plain-text line:
CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>model answer:
[
{
"tool": "create_ticket",
"args": {
"title": "dashboard shows stale numbers",
"priority": 9
}
},
{
"tool": "escalate",
"args": {
"ticket_id": "TCK-1"
}
},
{
"tool": "assign",
"args": {
"ticket_id": "TCK-1",
"agent": "novak"
}
},
{
"tool": "create_ticket",
"args": {
"title": "invoice total wrong",
"priority": 5
}
},
{
"tool": "assign",
"args": {
"ticket_id": "TCK-2",
"agent": "chen"
}
},
{
"tool": "create_ticket",
"args": {
"title": "invoice total wrong",
"priority": 5
}
},
{
"tool": "close_ticket",
"args": {
"ticket_id": "TCK-3",
"resolution": "duplicate of TCK-2"
}
},
{
"tool": "create_ticket",
"args": {
"title": "dashboard shows stale numbers",
"priority": 9
}
},
{
"tool": "close_ticket",
"args": {
"ticket_id": "TCK-4",
"resolution": "duplicate of TCK-1"
}
}
]wrongagentic.tools.deploy-v1conf — · 8.1s · $0.000 · 0 tok
question
You are a release agent operating strictly through tool calls.
TOOL CATALOG:
- deploy(service: string) — deploys a service; FAILS if any dependency is not yet deployed
- health_check(service: string) — runs a post-deploy health probe
- rollback(service: string) — reverts a bad deploy (nothing fails in this scenario)
- scale(service: string, replicas: int) — changes replica count (out of scope)
- restart(service: string) — restarts a service (out of scope)
SERVICES AND THEIR DEPENDENCIES (a service can only be deployed after ALL its dependencies):
- search: billing
- notifier: billing
- billing: (none)
- reports: billing
POLICY:
- Deploy in waves: in each wave, deploy every service whose dependencies are already deployed, in alphabetical order; repeat until all services are deployed.
- The service "notifier" is flagged unstable: call health_check on it immediately after deploying it.
- Call nothing else.
Output the full ordered sequence of tool calls needed to accomplish the goal, as a JSON array of objects {"tool": string, "args": object}. Use exactly the tool and argument names from the catalog. Do not call any tool that is not required.
Give your final answer inside ONE fenced code block (```), containing exactly the required content and nothing else. After the code block, add one plain-text line:
CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>model answer:
(none extracted)correctagentic.tools.ledger-v1conf 100% · 1.9s · $0.015 · 1623 tok
question
You are a treasury agent operating strictly through tool calls.
TOOL CATALOG:
- transfer(from: string, to: string, amount: int) — moves funds between accounts; FAILS if it would overdraw the source
- top_up_from_reserve(account: string, amount: int) — adds funds to an account from the corporate reserve
- get_balance(account: string) — reads a balance (you already have all balances below — reads are unnecessary and forbidden)
- freeze_account(account: string) — compliance freeze (not part of this task)
- convert_currency(account: string, currency: string) — FX conversion (all amounts are already in USD)
OPENING BALANCES:
- bravo: $885
- alpha: $131
- lima: $685
PAYMENT INSTRUCTIONS (execute in exactly this order):
1. pay $265 from "bravo" to "alpha"
2. pay $390 from "alpha" to "bravo"
3. pay $337 from "lima" to "bravo"
4. pay $449 from "bravo" to "alpha"
RULES:
- transfer() fails on overdraft. If a payment would overdraw its source account at the moment of execution, first call top_up_from_reserve() on the source with EXACTLY the shortfall (no more, no less), then execute the transfer.
- Track balances as they change: earlier payments affect later ones.
- Never call tools that are not needed.
Output the full ordered sequence of tool calls needed to accomplish the goal, as a JSON array of objects {"tool": string, "args": object}. Use exactly the tool and argument names from the catalog. Do not call any tool that is not required.
Give your final answer inside ONE fenced code block (```), containing exactly the required content and nothing else. After the code block, add one plain-text line:
CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>model answer:
[
{
"tool": "transfer",
"args": {
"from": "bravo",
"to": "alpha",
"amount": 265
}
},
{
"tool": "transfer",
"args": {
"from": "alpha",
"to": "bravo",
"amount": 390
}
},
{
"tool": "transfer",
"args": {
"from": "lima",
"to": "bravo",
"amount": 337
}
},
{
"tool": "transfer",
"args": {
"from": "bravo",
"to": "alpha",
"amount": 449
}
}
]correctagentic.tools.context-load-v1conf 100% · 2.0s · $0.036 · 3645 tok
question
You are an order-operations agent working strictly through tool calls.
TOOL CATALOG:
- restock(item: string, qty: int) — reorders stock for a large pending order
- cancel_order(order_id: int) — cancels a small pending order
- ship_order(order_id: int) — ships a paid order (out of scope here)
- refund(order_id: int, amount: int) — refunds a customer (out of scope here)
- notify_customer(customer: string, message: string) — sends a notification (not required by this policy)
ORDER LEDGER (121 records, format: id|customer|region|item|qty|status):
```
1653|gale|west|pump|21|paid
1733|acme|north|gasket|37|paid
1787|dorian|west|valve|43|held
1670|gale|east|pump|48|pending
1384|dorian|north|pump|45|paid
1677|gale|east|rotor|54|shipped
1331|birch|west|gasket|60|shipped
1424|ionic|north|cable|11|shipped
1558|harbor|south|cable|17|shipped
1742|dorian|west|frame|64|paid
1553|gale|north|panel|58|held
1744|cobalt|east|rotor|48|pending
1480|ionic|west|sensor|22|paid
1425|ember|west|frame|87|held
1697|gale|north|rotor|18|paid
1458|gale|east|gasket|83|shipped
1434|cobalt|north|gasket|29|held
1735|ember|east|frame|82|paid
1520|birch|east|cable|83|held
1324|birch|east|rotor|65|pending
1463|juno|south|valve|80|paid
1400|juno|east|frame|54|pending
1663|ember|east|cable|65|shipped
1778|birch|east|pump|35|shipped
1539|ember|east|rotor|91|paid
1315|birch|west|valve|19|pending
1405|juno|west|cable|53|pending
1363|ionic|north|gasket|85|pending
1642|acme|east|gasket|35|pending
1440|gale|south|sensor|71|paid
1413|dorian|north|pump|99|paid
1617|gale|east|gasket|95|pending
1370|gale|north|valve|92|paid
1560|ionic|south|panel|32|held
1389|ember|north|frame|93|paid
1664|cobalt|east|frame|37|held
1497|ionic|east|valve|16|paid
1656|juno|east|sensor|51|pending
1563|gale|north|frame|72|pending
1761|cobalt|south|frame|13|held
1746|birch|south|cable|87|pending
1410|ember|north|pump|57|shipped
1759|ember|east|valve|55|shipped
1449|ionic|west|panel|73|pending
1374|ember|north|rotor|31|shipped
1504|birch|east|cable|39|held
1478|gale|west|sensor|94|held
1517|ember|south|pump|16|shipped
1579|cobalt|east|cable|61|shipped
1352|dorian|south|sensor|27|shipped
1447|gale|south|frame|48|held
1651|fulton|north|pump|70|held
1709|ionic|east|sensor|25|pending
1417|ionic|south|sensor|88|pending
1730|gale|north|cable|11|pending
1455|juno|west|sensor|25|pending
1703|birch|north|rotor|42|paid
1639|gale|south|frame|94|shipped
1738|ionic|north|rotor|28|pending
1793|juno|south|rotor|41|pending
1695|fulton|south|sensor|74|paid
1719|acme|north|gasket|15|held
1613|acme|west|cable|96|shipped
1506|cobalt|south|rotor|50|held
1634|harbor|west|frame|53|paid
1756|fulton|south|valve|77|paid
1784|fulton|north|valve|72|shipped
1345|fulton|west|valve|38|paid
1427|acme|north|pump|59|pending
1348|harbor|west|sensor|43|paid
1322|birch|west|rotor|64|pending
1472|harbor|south|rotor|17|paid
1627|ionic|north|sensor|58|paid
1412|ember|east|gasket|66|shipped
1686|ionic|south|gasket|21|pending
1530|dorian|west|pump|39|pending
1562|fulton|south|sensor|24|shipped
1319|birch|north|valve|21|pending
1487|juno|east|frame|92|paid
1335|birch|west|panel|37|pending
1321|birch|west|pump|43|paid
1615|birch|east|panel|47|shipped
1624|ember|south|frame|91|shipped
1716|juno|west|panel|88|paid
1649|fulton|east|rotor|69|paid
1515|juno|west|valve|61|shipped
1490|dorian|west|cable|77|shipped
1774|harbor|south|panel|83|held
1358|acme|west|gasket|43|held
1595|dorian|west|valve|60|shipped
1692|gale|south|valve|55|shipped
1575|ionic|south|panel|37|held
1568|dorian|south|pump|36|pending
1760|cobalt|south|rotor|54|paid
1608|ember|north|pump|41|pending
1470|ember|west|panel|96|pending
1603|birch|west|gasket|66|held
1751|ionic|east|cable|46|paid
1573|harbor|east|valve|61|pending
1511|ionic|west|sensor|66|pending
1542|ember|north|valve|16|shipped
1665|gale|west|pump|99|paid
1586|cobalt|south|gasket|51|paid
1377|dorian|east|pump|28|paid
1590|birch|south|pump|36|paid
1584|acme|south|frame|37|shipped
1420|dorian|south|pump|62|paid
1396|harbor|north|valve|66|paid
1546|gale|south|valve|13|paid
1336|birch|east|sensor|98|pending
1570|birch|north|pump|39|held
1600|dorian|west|sensor|59|shipped
1390|birch|south|cable|56|pending
1682|birch|south|sensor|64|paid
1550|birch|east|pump|27|pending
1338|birch|west|gasket|18|shipped
1724|acme|west|cable|91|shipped
1768|fulton|north|pump|47|paid
1536|acme|north|rotor|96|pending
1610|ionic|south|panel|23|held
1526|gale|north|cable|39|paid
```
POLICY (apply exactly):
- Consider ONLY orders where customer = "birch" AND region = "west" AND status = "pending".
- Process those orders in ASCENDING order id.
- For each: if qty > 51, call restock(item, qty) with that order's item and qty; otherwise call cancel_order(order_id).
- Call nothing else. Every other record is irrelevant no matter how similar it looks.
Output the full ordered sequence of tool calls as a JSON array of {"tool": string, "args": object}.
Give your final answer inside ONE fenced code block (```), containing exactly the required content and nothing else. After the code block, add one plain-text line:
CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>model answer:
[
{"tool": "cancel_order", "args": {"order_id": 1315}},
{"tool": "restock", "args": {"item": "rotor", "qty": 64}},
{"tool": "cancel_order", "args": {"order_id": 1335}}
]correctagentic.tools.triage-v1conf 100% · 2.3s · $0.016 · 1645 tok
question
You operate a support desk strictly through tool calls.
TOOL CATALOG:
- create_ticket(title: string, priority: int) — opens a ticket; IDs are assigned sequentially: the 1st created ticket is "TCK-1", the 2nd "TCK-2", etc.
- assign(ticket_id: string, agent: string) — assigns an open ticket to an agent
- escalate(ticket_id: string) — marks a ticket as escalated
- close_ticket(ticket_id: string, resolution: string) — closes a ticket with a resolution note
- send_email(to: string, body: string) — sends an email (not part of the triage policy)
- archive_ticket(ticket_id: string) — archives a closed ticket (nightly job does this automatically)
- set_reminder(ticket_id: string, hours: int) — sets a follow-up reminder
ROUTING POLICY (apply exactly, in this order, for each incident, processing incidents in the order listed):
1. Create a ticket for the incident (title = the incident text verbatim, priority as given).
2. If the incident is an exact duplicate of an earlier incident in this list, close its ticket immediately with resolution "duplicate of <ID of the earlier ticket>" and do nothing else for it.
3. Otherwise, if priority ≥ 9, escalate the ticket BEFORE assigning it.
4. Assign the ticket to the agent responsible for its category.
CATEGORY → AGENT:
- auth → rivera
- payments → haddad
- infra → tanaka
INCIDENTS:
1. "locked out after 2FA change" (category: auth, priority 8)
2. "locked out after 2FA change" (category: auth, priority 8)
3. "webhooks not delivered" (category: infra, priority 7)
Output the full ordered sequence of tool calls needed to accomplish the goal, as a JSON array of objects {"tool": string, "args": object}. Use exactly the tool and argument names from the catalog. Do not call any tool that is not required.
Give your final answer inside ONE fenced code block (```), containing exactly the required content and nothing else. After the code block, add one plain-text line:
CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>model answer:
[
{
"tool": "create_ticket",
"args": {
"title": "locked out after 2FA change",
"priority": 8
}
},
{
"tool": "assign",
"args": {
"ticket_id": "TCK-1",
"agent": "rivera"
}
},
{
"tool": "create_ticket",
"args": {
"title": "locked out after 2FA change",
"priority": 8
}
},
{
"tool": "close_ticket",
"args": {
"ticket_id": "TCK-2",
"resolution": "duplicate of TCK-1"
}
},
{
"tool": "create_ticket",
"args": {
"title": "webhooks not delivered",
"priority": 7
}
},
{
"tool": "assign",
"args": {
"ticket_id": "TCK-3",
"agent": "tanaka"
}
}
]correctagentic.tools.context-load-v1conf 100% · 1.8s · $0.074 · 7377 tok
question
You are an order-operations agent working strictly through tool calls.
TOOL CATALOG:
- restock(item: string, qty: int) — reorders stock for a large pending order
- cancel_order(order_id: int) — cancels a small pending order
- ship_order(order_id: int) — ships a paid order (out of scope here)
- refund(order_id: int, amount: int) — refunds a customer (out of scope here)
- notify_customer(customer: string, message: string) — sends a notification (not required by this policy)
ORDER LEDGER (261 records, format: id|customer|region|item|qty|status):
```
2210|ionic|west|sensor|12|shipped
1662|ionic|north|gasket|74|held
2073|juno|east|pump|81|pending
1427|cobalt|east|panel|14|pending
1869|dorian|west|rotor|67|shipped
1355|juno|north|rotor|45|pending
1432|harbor|east|panel|81|paid
1471|dorian|west|frame|28|pending
1915|gale|north|valve|36|paid
1307|gale|south|gasket|79|pending
1564|gale|north|valve|32|held
2223|birch|south|gasket|25|shipped
1451|fulton|north|valve|98|paid
1433|birch|east|rotor|51|pending
2150|birch|west|pump|49|paid
1880|birch|east|valve|14|shipped
1435|harbor|east|pump|76|pending
1693|birch|north|panel|18|held
1262|ember|north|panel|19|pending
2235|dorian|north|valve|70|pending
2108|dorian|south|pump|17|paid
1168|fulton|east|pump|19|pending
1600|gale|south|frame|78|shipped
1259|ionic|south|sensor|91|held
1926|gale|north|gasket|59|held
1895|birch|west|pump|38|held
2012|fulton|south|rotor|79|paid
1917|fulton|south|frame|66|held
2086|fulton|south|valve|52|paid
1977|cobalt|west|cable|62|pending
2080|acme|south|frame|33|shipped
1175|fulton|south|frame|39|pending
1206|fulton|south|sensor|27|pending
1352|birch|west|panel|95|pending
1884|ionic|west|gasket|30|paid
1191|fulton|west|sensor|98|pending
1992|birch|west|panel|94|pending
2110|gale|north|cable|42|held
1999|harbor|west|pump|54|shipped
1340|harbor|south|sensor|27|shipped
1272|cobalt|south|frame|81|held
1487|birch|south|rotor|49|pending
1528|harbor|east|frame|76|paid
1945|cobalt|north|gasket|39|shipped
1356|birch|east|panel|96|held
2211|harbor|west|cable|73|held
2009|fulton|east|sensor|66|pending
1916|dorian|south|rotor|58|pending
1655|cobalt|east|cable|73|pending
2107|cobalt|south|gasket|34|shipped
1860|fulton|north|sensor|43|shipped
1844|ember|south|valve|51|pending
1220|fulton|south|panel|97|pending
1745|dorian|north|panel|94|pending
1875|fulton|north|cable|93|paid
1652|fulton|east|sensor|62|pending
1372|ionic|east|sensor|89|paid
2094|cobalt|north|valve|77|shipped
1667|cobalt|west|frame|93|pending
1334|gale|east|panel|39|paid
1891|ionic|west|valve|80|held
2091|cobalt|north|gasket|17|paid
1421|ionic|south|cable|80|paid
1846|birch|north|valve|95|paid
1366|fulton|west|sensor|67|pending
1509|ionic|east|panel|50|held
1227|fulton|west|rotor|28|pending
2079|ember|west|cable|26|pending
1930|cobalt|south|rotor|83|paid
1444|acme|south|panel|98|shipped
1363|birch|north|sensor|94|shipped
1786|dorian|north|sensor|13|pending
1551|cobalt|east|gasket|26|held
1186|fulton|east|cable|39|pending
1416|juno|west|panel|99|paid
1212|fulton|east|frame|30|shipped
2015|dorian|east|frame|44|pending
1347|birch|north|frame|51|pending
2060|gale|east|frame|83|held
1622|birch|south|pump|92|pending
1295|harbor|east|gasket|48|shipped
1383|ionic|west|gasket|16|paid
2180|harbor|north|sensor|81|pending
1482|fulton|north|panel|23|paid
1338|ember|north|panel|20|shipped
1567|gale|north|rotor|72|held
2157|birch|south|frame|42|pending
1863|fulton|north|pump|98|paid
1811|juno|south|panel|64|pending
1801|juno|south|pump|74|pending
2025|juno|south|gasket|82|pending
1936|harbor|east|panel|71|held
1461|ionic|east|valve|91|pending
1615|ember|west|panel|10|pending
1380|acme|north|cable|65|held
1724|gale|south|pump|56|pending
1986|birch|north|cable|41|pending
1527|dorian|south|rotor|30|paid
1284|dorian|south|gasket|49|held
1822|fulton|west|sensor|35|pending
1469|fulton|east|pump|12|paid
1458|harbor|north|cable|55|paid
1752|acme|north|frame|78|held
1517|juno|west|rotor|31|shipped
1243|cobalt|east|panel|70|paid
1400|ionic|north|sensor|21|held
1790|cobalt|north|rotor|97|paid
2018|ionic|east|gasket|68|paid
2164|dorian|north|sensor|79|pending
1730|ember|west|frame|89|pending
1820|acme|north|gasket|75|pending
1477|cobalt|south|pump|25|held
2207|gale|east|pump|94|pending
2096|ionic|east|pump|47|pending
1679|birch|west|pump|37|shipped
1939|birch|south|gasket|37|held
2005|acme|south|frame|14|shipped
1904|fulton|west|pump|15|shipped
1388|birch|west|pump|15|shipped
1418|harbor|south|rotor|62|paid
2189|acme|west|valve|77|paid
1247|fulton|east|pump|83|shipped
1537|ionic|north|pump|92|shipped
1921|dorian|north|cable|62|pending
1942|birch|north|gasket|49|pending
1682|harbor|west|cable|59|held
2011|juno|south|cable|92|shipped
2230|cobalt|west|cable|45|shipped
2117|ionic|east|panel|15|pending
1588|cobalt|west|panel|40|pending
2044|gale|north|gasket|11|held
2170|fulton|east|frame|20|pending
2133|harbor|west|cable|15|shipped
2177|birch|west|gasket|36|shipped
2179|gale|west|rotor|33|held
1437|ember|east|gasket|16|shipped
2052|ember|north|gasket|90|pending
1680|fulton|east|panel|11|paid
2031|ember|south|panel|49|shipped
1689|cobalt|north|valve|23|held
1776|harbor|east|panel|46|pending
1290|dorian|north|valve|84|held
2229|ionic|east|sensor|56|paid
1732|cobalt|east|panel|25|paid
2176|harbor|west|pump|72|pending
1857|birch|west|cable|39|pending
2010|ember|east|pump|81|paid
1252|gale|west|pump|57|held
1269|harbor|south|gasket|79|shipped
1378|fulton|south|valve|39|held
1236|fulton|north|cable|93|held
2041|dorian|north|frame|18|shipped
1557|juno|north|cable|14|paid
1772|harbor|south|pump|28|held
1934|dorian|north|frame|62|pending
1407|fulton|east|frame|61|held
1502|ember|west|frame|75|paid
1621|acme|east|pump|32|shipped
2241|dorian|south|pump|78|pending
1712|ember|west|gasket|75|shipped
1941|birch|east|cable|63|shipped
1278|harbor|north|sensor|27|paid
2101|ember|south|frame|85|paid
1501|gale|west|cable|21|pending
2068|ember|north|valve|49|shipped
1221|fulton|east|gasket|29|shipped
2216|ember|east|sensor|79|held
1900|dorian|west|rotor|33|pending
1971|harbor|south|cable|91|pending
1694|cobalt|north|sensor|52|held
1581|cobalt|east|cable|66|pending
2201|gale|east|pump|84|shipped
1673|ionic|east|sensor|61|pending
1606|gale|south|cable|19|shipped
1488|ember|west|gasket|12|shipped
1219|fulton|east|panel|82|pending
1516|cobalt|south|rotor|22|paid
1718|gale|west|panel|13|paid
1231|fulton|east|valve|16|shipped
1833|acme|west|valve|76|held
1317|cobalt|west|frame|47|pending
2038|dorian|north|gasket|67|paid
1807|ionic|north|pump|63|held
1414|ionic|south|panel|13|pending
1638|birch|south|cable|86|paid
1964|ionic|north|sensor|80|paid
1739|juno|north|gasket|96|paid
1951|gale|south|rotor|24|held
1646|gale|north|frame|43|paid
1991|cobalt|east|panel|62|held
1520|dorian|west|gasket|80|pending
1967|gale|south|frame|80|shipped
1742|juno|west|pump|41|paid
2123|harbor|south|valve|32|shipped
2057|ember|south|sensor|82|pending
1765|gale|south|rotor|43|shipped
1222|fulton|east|pump|65|pending
2064|birch|north|valve|12|held
2195|acme|south|sensor|89|pending
1549|ember|south|cable|28|pending
1842|gale|west|rotor|33|pending
1982|gale|north|rotor|95|pending
2071|acme|east|frame|21|pending
1633|ember|south|valve|79|pending
1330|acme|west|gasket|38|paid
1874|ionic|east|frame|29|paid
1584|acme|south|rotor|20|shipped
2034|juno|south|pump|46|paid
1324|dorian|north|pump|43|shipped
1794|gale|east|sensor|87|paid
1535|ember|west|rotor|71|held
2114|dorian|east|frame|79|pending
1628|gale|east|panel|29|shipped
1699|acme|east|rotor|50|held
1570|fulton|south|pump|44|paid
1774|juno|south|gasket|71|held
1862|harbor|east|frame|79|shipped
1393|dorian|south|pump|90|paid
1198|fulton|east|panel|22|held
2145|ionic|south|sensor|84|pending
1983|fulton|west|sensor|65|pending
1758|birch|west|cable|54|paid
1345|ember|south|sensor|73|shipped
1283|ionic|south|cable|77|pending
1827|ember|north|rotor|90|paid
1495|birch|north|rotor|13|pending
1664|ionic|north|sensor|39|held
1574|juno|north|cable|91|held
2129|birch|west|pump|10|pending
1587|fulton|west|valve|32|pending
1911|ember|south|pump|89|shipped
1454|ionic|west|frame|40|held
1705|birch|north|valve|20|shipped
1958|cobalt|south|rotor|32|pending
1889|ionic|north|cable|52|pending
2240|dorian|east|cable|52|pending
1302|acme|north|valve|66|paid
1594|harbor|west|pump|67|shipped
2046|cobalt|south|frame|23|shipped
1726|ionic|west|sensor|89|shipped
1611|gale|west|cable|21|held
1425|acme|south|sensor|68|paid
1853|fulton|east|cable|31|held
1780|fulton|west|pump|59|shipped
1607|gale|east|pump|14|held
2237|ionic|east|frame|57|pending
2066|cobalt|west|valve|45|pending
1838|juno|west|gasket|35|pending
1301|dorian|south|cable|84|held
1310|gale|east|valve|93|held
1203|fulton|east|rotor|86|pending
1395|ionic|east|valve|70|paid
1543|fulton|east|rotor|43|shipped
2047|dorian|north|frame|76|held
1643|cobalt|west|sensor|94|paid
2186|cobalt|east|valve|66|held
1512|birch|east|panel|87|held
1467|birch|west|frame|65|held
1816|dorian|north|sensor|39|pending
2139|gale|west|rotor|26|held
1181|fulton|east|rotor|90|held
```
POLICY (apply exactly):
- Consider ONLY orders where customer = "fulton" AND region = "east" AND status = "pending".
- Process those orders in ASCENDING order id.
- For each: if qty > 54, call restock(item, qty) with that order's item and qty; otherwise call cancel_order(order_id).
- Call nothing else. Every other record is irrelevant no matter how similar it looks.
Output the full ordered sequence of tool calls as a JSON array of {"tool": string, "args": object}.
Give your final answer inside ONE fenced code block (```), containing exactly the required content and nothing else. After the code block, add one plain-text line:
CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>model answer:
[
{"tool": "cancel_order", "args": {"order_id": 1168}},
{"tool": "cancel_order", "args": {"order_id": 1186}},
{"tool": "restock", "args": {"item": "rotor", "qty": 86}},
{"tool": "restock", "args": {"item": "panel", "qty": 82}},
{"tool": "restock", "args": {"item": "pump", "qty": 65}},
{"tool": "restock", "args": {"item": "sensor", "qty": 62}},
{"tool": "restock", "args": {"item": "sensor", "qty": 66}},
{"tool": "cancel_order", "args": {"order_id": 2170}}
]correctagentic.tools.deploy-v1conf 100% · 1.9s · $0.013 · 1435 tok
question
You are a release agent operating strictly through tool calls.
TOOL CATALOG:
- deploy(service: string) — deploys a service; FAILS if any dependency is not yet deployed
- health_check(service: string) — runs a post-deploy health probe
- rollback(service: string) — reverts a bad deploy (nothing fails in this scenario)
- scale(service: string, replicas: int) — changes replica count (out of scope)
- restart(service: string) — restarts a service (out of scope)
SERVICES AND THEIR DEPENDENCIES (a service can only be deployed after ALL its dependencies):
- billing: reports
- reports: gateway
- gateway: (none)
- search: gateway
POLICY:
- Deploy in waves: in each wave, deploy every service whose dependencies are already deployed, in alphabetical order; repeat until all services are deployed.
- The service "search" is flagged unstable: call health_check on it immediately after deploying it.
- Call nothing else.
Output the full ordered sequence of tool calls needed to accomplish the goal, as a JSON array of objects {"tool": string, "args": object}. Use exactly the tool and argument names from the catalog. Do not call any tool that is not required.
Give your final answer inside ONE fenced code block (```), containing exactly the required content and nothing else. After the code block, add one plain-text line:
CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>model answer:
[
{
"tool": "deploy",
"args": {
"service": "gateway"
}
},
{
"tool": "deploy",
"args": {
"service": "reports"
}
},
{
"tool": "deploy",
"args": {
"service": "search"
}
},
{
"tool": "health_check",
"args": {
"service": "search"
}
},
{
"tool": "deploy",
"args": {
"service": "billing"
}
}
]correctagentic.tools.ledger-v1conf 100% · 1.6s · $0.018 · 1958 tok
question
You are a treasury agent operating strictly through tool calls.
TOOL CATALOG:
- transfer(from: string, to: string, amount: int) — moves funds between accounts; FAILS if it would overdraw the source
- top_up_from_reserve(account: string, amount: int) — adds funds to an account from the corporate reserve
- get_balance(account: string) — reads a balance (you already have all balances below — reads are unnecessary and forbidden)
- freeze_account(account: string) — compliance freeze (not part of this task)
- convert_currency(account: string, currency: string) — FX conversion (all amounts are already in USD)
OPENING BALANCES:
- bravo: $272
- oscar: $395
- alpha: $624
PAYMENT INSTRUCTIONS (execute in exactly this order):
1. pay $129 from "bravo" to "alpha"
2. pay $107 from "oscar" to "alpha"
3. pay $172 from "oscar" to "alpha"
4. pay $247 from "bravo" to "alpha"
RULES:
- transfer() fails on overdraft. If a payment would overdraw its source account at the moment of execution, first call top_up_from_reserve() on the source with EXACTLY the shortfall (no more, no less), then execute the transfer.
- Track balances as they change: earlier payments affect later ones.
- Never call tools that are not needed.
Output the full ordered sequence of tool calls needed to accomplish the goal, as a JSON array of objects {"tool": string, "args": object}. Use exactly the tool and argument names from the catalog. Do not call any tool that is not required.
Give your final answer inside ONE fenced code block (```), containing exactly the required content and nothing else. After the code block, add one plain-text line:
CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>model answer:
[
{
"tool": "transfer",
"args": {
"from": "bravo",
"to": "alpha",
"amount": 129
}
},
{
"tool": "transfer",
"args": {
"from": "oscar",
"to": "alpha",
"amount": 107
}
},
{
"tool": "transfer",
"args": {
"from": "oscar",
"to": "alpha",
"amount": 172
}
},
{
"tool": "top_up_from_reserve",
"args": {
"account": "bravo",
"amount": 104
}
},
{
"tool": "transfer",
"args": {
"from": "bravo",
"to": "alpha",
"amount": 247
}
}
]correctagentic.tools.triage-v1conf 100% · 1.8s · $0.014 · 1514 tok
question
You operate a support desk strictly through tool calls.
TOOL CATALOG:
- create_ticket(title: string, priority: int) — opens a ticket; IDs are assigned sequentially: the 1st created ticket is "TCK-1", the 2nd "TCK-2", etc.
- assign(ticket_id: string, agent: string) — assigns an open ticket to an agent
- escalate(ticket_id: string) — marks a ticket as escalated
- close_ticket(ticket_id: string, resolution: string) — closes a ticket with a resolution note
- send_email(to: string, body: string) — sends an email (not part of the triage policy)
- archive_ticket(ticket_id: string) — archives a closed ticket (nightly job does this automatically)
- set_reminder(ticket_id: string, hours: int) — sets a follow-up reminder
ROUTING POLICY (apply exactly, in this order, for each incident, processing incidents in the order listed):
1. Create a ticket for the incident (title = the incident text verbatim, priority as given).
2. If the incident is an exact duplicate of an earlier incident in this list, close its ticket immediately with resolution "duplicate of <ID of the earlier ticket>" and do nothing else for it.
3. Otherwise, if priority ≥ 8, escalate the ticket BEFORE assigning it.
4. Assign the ticket to the agent responsible for its category.
CATEGORY → AGENT:
- payments → tanaka
- infra → dubois
- auth → chen
INCIDENTS:
1. "refund double-charged" (category: payments, priority 8)
2. "refund double-charged" (category: payments, priority 8)
3. "SSO loop on login" (category: auth, priority 2)
Output the full ordered sequence of tool calls needed to accomplish the goal, as a JSON array of objects {"tool": string, "args": object}. Use exactly the tool and argument names from the catalog. Do not call any tool that is not required.
Give your final answer inside ONE fenced code block (```), containing exactly the required content and nothing else. After the code block, add one plain-text line:
CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>model answer:
[
{
"tool": "create_ticket",
"args": {
"title": "refund double-charged",
"priority": 8
}
},
{
"tool": "escalate",
"args": {
"ticket_id": "TCK-1"
}
},
{
"tool": "assign",
"args": {
"ticket_id": "TCK-1",
"agent": "tanaka"
}
},
{
"tool": "create_ticket",
"args": {
"title": "refund double-charged",
"priority": 8
}
},
{
"tool": "close_ticket",
"args": {
"ticket_id": "TCK-2",
"resolution": "duplicate of TCK-1"
}
},
{
"tool": "create_ticket",
"args": {
"title": "SSO loop on login",
"priority": 2
}
},
{
"tool": "assign",
"args": {
"ticket_id": "TCK-3",
"agent": "chen"
}
}
]correctagentic.tools.deploy-v1conf 100% · 8.0s · $0.012 · 1238 tok
question
You are a release agent operating strictly through tool calls.
TOOL CATALOG:
- deploy(service: string) — deploys a service; FAILS if any dependency is not yet deployed
- health_check(service: string) — runs a post-deploy health probe
- rollback(service: string) — reverts a bad deploy (nothing fails in this scenario)
- scale(service: string, replicas: int) — changes replica count (out of scope)
- restart(service: string) — restarts a service (out of scope)
SERVICES AND THEIR DEPENDENCIES (a service can only be deployed after ALL its dependencies):
- billing: auth-svc, search
- search: auth-svc
- reports: auth-svc, search
- auth-svc: (none)
POLICY:
- Deploy in waves: in each wave, deploy every service whose dependencies are already deployed, in alphabetical order; repeat until all services are deployed.
- The service "search" is flagged unstable: call health_check on it immediately after deploying it.
- Call nothing else.
Output the full ordered sequence of tool calls needed to accomplish the goal, as a JSON array of objects {"tool": string, "args": object}. Use exactly the tool and argument names from the catalog. Do not call any tool that is not required.
Give your final answer inside ONE fenced code block (```), containing exactly the required content and nothing else. After the code block, add one plain-text line:
CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>model answer:
[
{"tool": "deploy", "args": {"service": "auth-svc"}},
{"tool": "deploy", "args": {"service": "search"}},
{"tool": "health_check", "args": {"service": "search"}},
{"tool": "deploy", "args": {"service": "billing"}},
{"tool": "deploy", "args": {"service": "reports"}}
]correctagentic.tools.ledger-v1conf 100% · 10.0s · $0.017 · 1773 tok
question
You are a treasury agent operating strictly through tool calls.
TOOL CATALOG:
- transfer(from: string, to: string, amount: int) — moves funds between accounts; FAILS if it would overdraw the source
- top_up_from_reserve(account: string, amount: int) — adds funds to an account from the corporate reserve
- get_balance(account: string) — reads a balance (you already have all balances below — reads are unnecessary and forbidden)
- freeze_account(account: string) — compliance freeze (not part of this task)
- convert_currency(account: string, currency: string) — FX conversion (all amounts are already in USD)
OPENING BALANCES:
- kilo: $403
- bravo: $867
- lima: $462
PAYMENT INSTRUCTIONS (execute in exactly this order):
1. pay $225 from "lima" to "bravo"
2. pay $327 from "kilo" to "lima"
3. pay $483 from "lima" to "kilo"
4. pay $469 from "bravo" to "lima"
RULES:
- transfer() fails on overdraft. If a payment would overdraw its source account at the moment of execution, first call top_up_from_reserve() on the source with EXACTLY the shortfall (no more, no less), then execute the transfer.
- Track balances as they change: earlier payments affect later ones.
- Never call tools that are not needed.
Output the full ordered sequence of tool calls needed to accomplish the goal, as a JSON array of objects {"tool": string, "args": object}. Use exactly the tool and argument names from the catalog. Do not call any tool that is not required.
Give your final answer inside ONE fenced code block (```), containing exactly the required content and nothing else. After the code block, add one plain-text line:
CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>model answer:
[
{"tool": "transfer", "args": {"from": "lima", "to": "bravo", "amount": 225}},
{"tool": "transfer", "args": {"from": "kilo", "to": "lima", "amount": 327}},
{"tool": "transfer", "args": {"from": "lima", "to": "kilo", "amount": 483}},
{"tool": "transfer", "args": {"from": "bravo", "to": "lima", "amount": 469}}
]correctagentic.tools.context-load-v1conf 100% · 1.5s · $0.066 · 6633 tok
question
You are an order-operations agent working strictly through tool calls.
TOOL CATALOG:
- restock(item: string, qty: int) — reorders stock for a large pending order
- cancel_order(order_id: int) — cancels a small pending order
- ship_order(order_id: int) — ships a paid order (out of scope here)
- refund(order_id: int, amount: int) — refunds a customer (out of scope here)
- notify_customer(customer: string, message: string) — sends a notification (not required by this policy)
ORDER LEDGER (233 records, format: id|customer|region|item|qty|status):
```
1496|juno|south|cable|40|pending
1546|ionic|south|rotor|23|pending
1826|harbor|south|valve|57|shipped
2328|ember|south|valve|72|held
1875|ionic|north|cable|61|held
2286|ember|east|sensor|51|shipped
1894|harbor|south|pump|56|pending
1713|birch|west|cable|62|held
1504|juno|south|sensor|14|held
2016|gale|north|valve|83|paid
2313|dorian|north|pump|40|held
1863|birch|east|sensor|44|shipped
1963|acme|east|valve|45|paid
1472|juno|south|panel|23|held
1997|cobalt|east|panel|21|pending
1520|harbor|south|valve|68|paid
1590|dorian|north|frame|73|paid
2334|birch|south|cable|72|shipped
1912|dorian|east|valve|33|held
1579|harbor|south|sensor|71|paid
1945|dorian|east|gasket|55|paid
2179|dorian|west|sensor|63|paid
2222|harbor|east|cable|20|paid
2134|ember|south|pump|76|paid
1833|fulton|north|frame|65|held
2062|cobalt|south|frame|23|pending
1736|acme|west|valve|52|held
2076|juno|east|pump|90|held
1582|ionic|east|cable|12|held
1691|juno|west|valve|51|held
2191|harbor|north|valve|73|held
1744|fulton|south|cable|89|paid
1720|fulton|west|cable|18|paid
2206|ember|north|pump|77|held
2102|birch|west|sensor|81|held
2030|harbor|south|rotor|12|pending
1693|acme|east|frame|53|paid
1621|acme|east|pump|75|paid
1521|fulton|east|frame|81|held
1663|fulton|east|pump|95|shipped
1585|gale|south|valve|41|pending
2160|acme|west|frame|17|pending
1968|dorian|west|rotor|59|paid
2111|acme|east|valve|89|held
1670|cobalt|south|gasket|69|paid
1694|harbor|east|valve|50|pending
1779|ember|north|valve|47|paid
2333|dorian|south|pump|45|shipped
1791|gale|west|panel|96|paid
1651|juno|east|cable|77|shipped
2281|gale|north|gasket|61|pending
2175|harbor|north|panel|50|held
1500|juno|west|frame|10|pending
1814|ember|south|cable|24|held
2228|ionic|west|valve|16|shipped
1614|fulton|south|gasket|11|held
1487|juno|south|rotor|16|held
1811|cobalt|south|frame|97|pending
2105|fulton|west|gasket|26|held
1563|cobalt|west|panel|84|pending
1484|juno|west|frame|54|pending
1823|fulton|west|rotor|74|held
2047|acme|south|valve|97|shipped
2040|acme|east|rotor|64|paid
2100|acme|east|frame|35|held
2000|birch|north|sensor|52|shipped
2184|dorian|south|frame|77|pending
2383|dorian|south|valve|57|shipped
2268|ember|east|rotor|44|pending
1908|ember|east|rotor|16|pending
2069|ember|north|sensor|97|held
1729|fulton|north|gasket|79|pending
2276|ember|east|frame|82|pending
1601|ionic|south|valve|76|shipped
2128|ionic|north|sensor|31|shipped
1794|acme|north|pump|19|paid
1938|acme|north|gasket|23|held
1698|ionic|north|frame|60|shipped
1634|gale|north|gasket|37|shipped
1913|gale|north|pump|23|held
2181|birch|west|panel|70|shipped
1731|acme|west|frame|13|paid
1926|juno|east|sensor|40|paid
1506|birch|west|frame|64|held
1465|juno|south|pump|38|pending
2269|acme|east|sensor|83|pending
2216|ionic|north|cable|21|paid
1801|ember|south|pump|97|shipped
1737|birch|west|panel|94|shipped
1685|cobalt|south|cable|14|held
1542|harbor|north|pump|74|held
1855|birch|east|cable|27|held
1991|fulton|west|gasket|56|shipped
2088|cobalt|west|pump|22|held
1919|fulton|east|cable|70|pending
1981|ionic|west|pump|48|pending
1666|juno|west|valve|98|paid
2157|ionic|south|frame|80|paid
2117|acme|east|gasket|48|held
2090|fulton|west|valve|91|paid
2368|dorian|east|sensor|61|shipped
2209|dorian|south|gasket|83|pending
1806|acme|west|panel|22|held
1949|ember|east|valve|33|paid
1933|cobalt|east|valve|27|held
1726|ionic|south|sensor|35|held
1707|cobalt|east|pump|34|pending
2145|harbor|east|sensor|13|pending
2178|acme|east|rotor|24|paid
2311|acme|south|rotor|90|shipped
2365|ionic|east|rotor|28|pending
1509|birch|west|frame|15|paid
1849|ember|north|rotor|67|held
1478|juno|south|cable|99|pending
1529|gale|north|sensor|93|held
2022|acme|east|rotor|10|paid
2151|fulton|south|valve|43|pending
1682|ionic|east|sensor|62|pending
1969|harbor|south|pump|80|pending
2095|ionic|west|cable|88|paid
1535|gale|south|sensor|12|pending
1750|fulton|east|cable|57|held
2387|birch|north|valve|89|shipped
1816|ionic|west|gasket|90|paid
2150|dorian|west|pump|52|pending
1552|ionic|west|gasket|72|paid
2394|juno|east|rotor|36|shipped
2254|gale|south|sensor|71|held
1835|acme|east|cable|89|held
2182|cobalt|south|frame|94|paid
1656|ionic|west|pump|43|shipped
2224|gale|east|sensor|81|shipped
1869|fulton|west|valve|15|held
1514|cobalt|south|panel|52|shipped
1541|dorian|east|sensor|45|held
2053|ember|south|sensor|42|held
2162|ionic|east|gasket|43|paid
1551|dorian|west|panel|89|pending
2144|cobalt|south|valve|77|held
2379|harbor|north|valve|50|pending
1597|acme|south|rotor|66|paid
1558|harbor|west|frame|29|held
1756|fulton|south|frame|76|held
2198|juno|south|frame|90|held
1783|juno|south|rotor|70|pending
1956|acme|east|sensor|85|pending
2055|juno|south|rotor|35|paid
1862|acme|north|frame|33|held
2225|ionic|west|pump|96|shipped
2024|fulton|west|frame|50|shipped
1549|birch|west|valve|60|held
1977|cobalt|south|pump|58|pending
1493|juno|east|cable|45|pending
2261|dorian|south|gasket|40|paid
2196|birch|east|rotor|12|shipped
2123|harbor|south|valve|31|pending
2006|dorian|east|rotor|95|paid
2031|juno|south|gasket|53|paid
1877|ionic|north|gasket|31|held
1540|fulton|south|rotor|74|pending
2230|fulton|north|frame|48|shipped
2258|birch|east|pump|25|paid
1677|dorian|east|gasket|59|pending
2307|ionic|east|panel|97|shipped
1467|juno|north|frame|78|pending
2354|dorian|north|cable|22|held
1645|acme|north|pump|20|pending
1505|acme|north|cable|17|shipped
1607|juno|south|panel|69|held
2208|acme|east|sensor|67|shipped
2233|birch|west|cable|75|paid
1772|acme|north|gasket|75|pending
2389|harbor|south|valve|59|paid
1970|ionic|north|frame|41|held
2131|birch|south|pump|16|held
1640|fulton|west|sensor|22|paid
2227|ember|west|valve|11|shipped
1632|dorian|east|cable|14|paid
2071|fulton|east|sensor|76|pending
2238|fulton|north|frame|23|held
1892|dorian|north|valve|11|shipped
2161|juno|west|rotor|46|held
1526|harbor|east|gasket|80|pending
1763|cobalt|south|sensor|41|shipped
1899|acme|south|frame|59|paid
2288|birch|west|pump|84|pending
2048|cobalt|south|sensor|19|shipped
1489|juno|south|sensor|96|pending
2169|birch|north|frame|22|paid
1767|ember|south|pump|48|held
1842|birch|east|panel|84|paid
1986|dorian|north|sensor|84|shipped
1885|juno|east|gasket|53|pending
2244|juno|south|sensor|50|shipped
2316|acme|west|pump|62|held
2301|birch|west|gasket|80|paid
1911|fulton|north|gasket|93|paid
1880|ember|north|rotor|92|shipped
1494|juno|south|panel|61|shipped
2274|ionic|west|pump|97|shipped
1572|gale|north|sensor|76|paid
1734|ionic|east|sensor|38|shipped
2361|fulton|north|cable|54|paid
2204|birch|south|gasket|78|pending
1966|gale|south|sensor|84|shipped
2374|birch|east|rotor|35|held
2154|birch|south|frame|15|pending
2294|ember|west|pump|81|shipped
1704|cobalt|north|rotor|99|shipped
2248|gale|west|pump|31|held
2012|fulton|east|pump|40|paid
2348|ember|south|panel|91|paid
2323|gale|north|sensor|19|shipped
1888|gale|south|rotor|93|paid
2084|dorian|south|panel|68|pending
1569|juno|south|pump|31|paid
1776|ember|south|valve|87|shipped
1909|birch|north|frame|97|held
2020|cobalt|east|frame|78|held
2372|juno|west|rotor|54|paid
1636|cobalt|south|sensor|76|paid
1530|juno|west|rotor|92|pending
1759|acme|south|gasket|26|paid
2341|fulton|west|pump|89|pending
1627|fulton|west|rotor|14|held
2141|gale|west|cable|80|held
1903|juno|west|valve|39|held
2077|gale|north|pump|59|shipped
2033|acme|east|gasket|64|held
2304|dorian|west|valve|33|pending
1785|dorian|south|panel|64|paid
2140|juno|north|panel|60|shipped
1976|ember|south|pump|93|shipped
```
POLICY (apply exactly):
- Consider ONLY orders where customer = "juno" AND region = "south" AND status = "pending".
- Process those orders in ASCENDING order id.
- For each: if qty > 48, call restock(item, qty) with that order's item and qty; otherwise call cancel_order(order_id).
- Call nothing else. Every other record is irrelevant no matter how similar it looks.
Output the full ordered sequence of tool calls as a JSON array of {"tool": string, "args": object}.
Give your final answer inside ONE fenced code block (```), containing exactly the required content and nothing else. After the code block, add one plain-text line:
CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>model answer:
[
{"tool": "cancel_order", "args": {"order_id": 1465}},
{"tool": "restock", "args": {"item": "cable", "qty": 99}},
{"tool": "restock", "args": {"item": "sensor", "qty": 96}},
{"tool": "cancel_order", "args": {"order_id": 1496}},
{"tool": "restock", "args": {"item": "rotor", "qty": 70}}
]wrongagentic.tools.triage-v1conf 100% · 2.3s · $0.018 · 1941 tok
question
You operate a support desk strictly through tool calls.
TOOL CATALOG:
- create_ticket(title: string, priority: int) — opens a ticket; IDs are assigned sequentially: the 1st created ticket is "TCK-1", the 2nd "TCK-2", etc.
- assign(ticket_id: string, agent: string) — assigns an open ticket to an agent
- escalate(ticket_id: string) — marks a ticket as escalated
- close_ticket(ticket_id: string, resolution: string) — closes a ticket with a resolution note
- send_email(to: string, body: string) — sends an email (not part of the triage policy)
- archive_ticket(ticket_id: string) — archives a closed ticket (nightly job does this automatically)
- set_reminder(ticket_id: string, hours: int) — sets a follow-up reminder
ROUTING POLICY (apply exactly, in this order, for each incident, processing incidents in the order listed):
1. Create a ticket for the incident (title = the incident text verbatim, priority as given).
2. If the incident is an exact duplicate of an earlier incident in this list, close its ticket immediately with resolution "duplicate of <ID of the earlier ticket>" and do nothing else for it.
3. Otherwise, if priority ≥ 7, escalate the ticket BEFORE assigning it.
4. Assign the ticket to the agent responsible for its category.
CATEGORY → AGENT:
- payments → tanaka
- auth → haddad
- data → novak
INCIDENTS:
1. "invoice total wrong" (category: payments, priority 6)
2. "invoice total wrong" (category: payments, priority 6)
3. "dashboard shows stale numbers" (category: data, priority 9)
4. "invoice total wrong" (category: payments, priority 8)
Output the full ordered sequence of tool calls needed to accomplish the goal, as a JSON array of objects {"tool": string, "args": object}. Use exactly the tool and argument names from the catalog. Do not call any tool that is not required.
Give your final answer inside ONE fenced code block (```), containing exactly the required content and nothing else. After the code block, add one plain-text line:
CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>model answer:
[
{
"tool": "create_ticket",
"args": {
"title": "invoice total wrong",
"priority": 6
}
},
{
"tool": "assign",
"args": {
"ticket_id": "TCK-1",
"agent": "tanaka"
}
},
{
"tool": "create_ticket",
"args": {
"title": "invoice total wrong",
"priority": 6
}
},
{
"tool": "close_ticket",
"args": {
"ticket_id": "TCK-2",
"resolution": "duplicate of TCK-1"
}
},
{
"tool": "create_ticket",
"args": {
"title": "dashboard shows stale numbers",
"priority": 9
}
},
{
"tool": "escalate",
"args": {
"ticket_id": "TCK-3"
}
},
{
"tool": "assign",
"args": {
"ticket_id": "TCK-3",
"agent": "novak"
}
},
{
"tool": "create_ticket",
"args": {
"title": "invoice total wrong",
"priority": 8
}
},
{
"tool": "close_ticket",
"args": {
"ticket_id": "TCK-4",
"resolution": "duplicate of TCK-1"
}
}
]correctagentic.tools.deploy-v1conf 100% · 1.5s · $0.014 · 1497 tok
question
You are a release agent operating strictly through tool calls.
TOOL CATALOG:
- deploy(service: string) — deploys a service; FAILS if any dependency is not yet deployed
- health_check(service: string) — runs a post-deploy health probe
- rollback(service: string) — reverts a bad deploy (nothing fails in this scenario)
- scale(service: string, replicas: int) — changes replica count (out of scope)
- restart(service: string) — restarts a service (out of scope)
SERVICES AND THEIR DEPENDENCIES (a service can only be deployed after ALL its dependencies):
- search: auth-svc, notifier
- notifier: (none)
- reports: notifier
- auth-svc: notifier
POLICY:
- Deploy in waves: in each wave, deploy every service whose dependencies are already deployed, in alphabetical order; repeat until all services are deployed.
- The service "notifier" is flagged unstable: call health_check on it immediately after deploying it.
- Call nothing else.
Output the full ordered sequence of tool calls needed to accomplish the goal, as a JSON array of objects {"tool": string, "args": object}. Use exactly the tool and argument names from the catalog. Do not call any tool that is not required.
Give your final answer inside ONE fenced code block (```), containing exactly the required content and nothing else. After the code block, add one plain-text line:
CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>model answer:
[
{
"tool": "deploy",
"args": {
"service": "notifier"
}
},
{
"tool": "health_check",
"args": {
"service": "notifier"
}
},
{
"tool": "deploy",
"args": {
"service": "auth-svc"
}
},
{
"tool": "deploy",
"args": {
"service": "reports"
}
},
{
"tool": "deploy",
"args": {
"service": "search"
}
}
]correctagentic.tools.ledger-v1conf 100% · 7.8s · $0.025 · 2655 tok
question
You are a treasury agent operating strictly through tool calls.
TOOL CATALOG:
- transfer(from: string, to: string, amount: int) — moves funds between accounts; FAILS if it would overdraw the source
- top_up_from_reserve(account: string, amount: int) — adds funds to an account from the corporate reserve
- get_balance(account: string) — reads a balance (you already have all balances below — reads are unnecessary and forbidden)
- freeze_account(account: string) — compliance freeze (not part of this task)
- convert_currency(account: string, currency: string) — FX conversion (all amounts are already in USD)
OPENING BALANCES:
- bravo: $566
- tango: $134
- delta: $808
PAYMENT INSTRUCTIONS (execute in exactly this order):
1. pay $271 from "delta" to "tango"
2. pay $599 from "delta" to "tango"
3. pay $216 from "tango" to "delta"
4. pay $343 from "delta" to "bravo"
RULES:
- transfer() fails on overdraft. If a payment would overdraw its source account at the moment of execution, first call top_up_from_reserve() on the source with EXACTLY the shortfall (no more, no less), then execute the transfer.
- Track balances as they change: earlier payments affect later ones.
- Never call tools that are not needed.
Output the full ordered sequence of tool calls needed to accomplish the goal, as a JSON array of objects {"tool": string, "args": object}. Use exactly the tool and argument names from the catalog. Do not call any tool that is not required.
Give your final answer inside ONE fenced code block (```), containing exactly the required content and nothing else. After the code block, add one plain-text line:
CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>model answer:
[
{
"tool": "transfer",
"args": {
"from": "delta",
"to": "tango",
"amount": 271
}
},
{
"tool": "top_up_from_reserve",
"args": {
"account": "delta",
"amount": 62
}
},
{
"tool": "transfer",
"args": {
"from": "delta",
"to": "tango",
"amount": 599
}
},
{
"tool": "transfer",
"args": {
"from": "tango",
"to": "delta",
"amount": 216
}
},
{
"tool": "top_up_from_reserve",
"args": {
"account": "delta",
"amount": 127
}
},
{
"tool": "transfer",
"args": {
"from": "delta",
"to": "bravo",
"amount": 343
}
}
]correctagentic.tools.context-load-v1conf 100% · 1.7s · $0.062 · 6332 tok
question
You are an order-operations agent working strictly through tool calls.
TOOL CATALOG:
- restock(item: string, qty: int) — reorders stock for a large pending order
- cancel_order(order_id: int) — cancels a small pending order
- ship_order(order_id: int) — ships a paid order (out of scope here)
- refund(order_id: int, amount: int) — refunds a customer (out of scope here)
- notify_customer(customer: string, message: string) — sends a notification (not required by this policy)
ORDER LEDGER (194 records, format: id|customer|region|item|qty|status):
```
1518|gale|south|valve|85|paid
1552|juno|east|sensor|20|shipped
1488|birch|south|valve|60|pending
1983|fulton|south|gasket|18|pending
1967|ember|west|sensor|67|pending
1952|ember|south|panel|44|shipped
1595|dorian|east|sensor|94|held
1739|birch|west|cable|42|paid
1746|juno|south|frame|97|paid
2000|birch|south|valve|83|pending
1436|dorian|east|cable|21|pending
1843|harbor|west|frame|18|pending
1757|cobalt|north|valve|10|shipped
1316|cobalt|east|gasket|59|pending
1921|ionic|east|panel|54|shipped
1768|birch|south|pump|24|shipped
1312|cobalt|south|sensor|71|pending
1873|gale|east|frame|23|held
1969|juno|south|panel|55|pending
1639|birch|north|sensor|82|paid
1785|harbor|north|pump|16|paid
1591|birch|north|sensor|91|pending
1691|cobalt|north|panel|96|shipped
1461|ember|west|valve|24|paid
1896|ionic|north|panel|32|pending
1787|gale|east|panel|86|shipped
1623|gale|west|gasket|96|shipped
1938|cobalt|north|sensor|67|held
1821|birch|south|rotor|20|shipped
1440|ember|south|panel|85|pending
1741|birch|south|cable|51|shipped
1600|ionic|east|sensor|12|pending
1352|harbor|west|panel|42|paid
1706|cobalt|south|cable|97|shipped
1998|birch|south|rotor|63|paid
1740|dorian|south|panel|69|pending
1645|cobalt|east|pump|23|paid
1588|dorian|south|cable|19|pending
1547|harbor|west|frame|83|pending
1689|dorian|south|panel|97|held
1425|cobalt|south|cable|52|shipped
2007|fulton|north|panel|42|held
1505|ionic|west|pump|49|held
1522|cobalt|north|gasket|83|held
1287|cobalt|south|panel|48|pending
1646|cobalt|north|frame|94|held
1723|cobalt|south|pump|20|paid
1331|cobalt|west|pump|14|paid
1512|harbor|west|rotor|48|shipped
1413|juno|south|frame|39|shipped
1451|gale|south|gasket|63|pending
1933|harbor|east|gasket|39|held
1403|cobalt|south|rotor|15|shipped
1504|cobalt|north|frame|39|paid
1752|ember|east|gasket|78|paid
1885|gale|north|sensor|77|held
1677|ionic|south|pump|53|pending
1978|dorian|north|gasket|85|paid
1931|juno|south|panel|62|paid
1574|harbor|south|panel|10|held
1657|ionic|east|valve|45|paid
1734|dorian|east|frame|36|pending
1912|gale|west|gasket|42|paid
1690|harbor|west|panel|32|pending
1406|juno|west|gasket|72|shipped
1289|cobalt|north|rotor|48|pending
1474|dorian|south|cable|27|shipped
1557|harbor|west|sensor|27|held
1584|juno|south|cable|88|paid
1844|juno|south|frame|98|pending
1697|fulton|east|rotor|41|held
1986|juno|south|pump|73|pending
1981|ionic|south|valve|93|pending
1322|cobalt|south|sensor|44|shipped
1860|fulton|west|frame|87|held
1271|cobalt|south|sensor|41|pending
1800|juno|north|cable|86|pending
1548|juno|north|pump|24|paid
1711|birch|west|panel|85|pending
2034|birch|south|rotor|61|paid
1960|birch|north|rotor|57|held
1909|dorian|south|pump|12|held
1598|cobalt|south|valve|32|pending
1433|dorian|west|sensor|49|paid
1278|cobalt|west|valve|24|pending
1320|cobalt|east|sensor|41|pending
1458|dorian|west|gasket|16|paid
1782|cobalt|north|panel|77|held
1443|birch|north|rotor|38|held
1891|dorian|west|valve|70|held
1904|juno|east|gasket|59|pending
1749|gale|north|gasket|97|held
1908|birch|east|sensor|63|pending
1418|fulton|west|panel|29|held
1881|gale|east|gasket|39|pending
1944|cobalt|west|sensor|62|shipped
1720|gale|west|panel|40|paid
1306|cobalt|south|rotor|25|shipped
1405|birch|west|valve|47|pending
1762|ember|east|valve|93|shipped
1837|ember|north|panel|96|held
1544|juno|south|sensor|39|shipped
1347|ionic|south|rotor|70|held
1297|cobalt|south|gasket|62|pending
1992|cobalt|south|rotor|17|shipped
1827|acme|east|gasket|40|held
1888|birch|north|pump|57|shipped
1669|birch|south|frame|17|paid
1849|juno|west|frame|52|paid
1564|birch|east|pump|24|pending
1394|ionic|south|gasket|36|pending
1672|fulton|north|cable|26|paid
1435|harbor|east|sensor|82|shipped
1996|ember|north|gasket|94|paid
1775|dorian|south|sensor|23|held
1615|dorian|east|pump|14|held
1388|birch|south|panel|15|shipped
1317|cobalt|south|cable|59|shipped
1807|ember|east|gasket|68|pending
1430|gale|south|pump|29|paid
1485|gale|north|valve|76|paid
1582|acme|west|cable|78|pending
1328|juno|west|panel|50|paid
1633|ember|west|sensor|56|held
1630|ionic|north|pump|44|paid
1728|ember|south|frame|32|paid
2020|ember|south|valve|33|pending
1358|acme|south|gasket|88|paid
1419|ember|west|panel|25|shipped
1463|dorian|east|gasket|59|paid
1713|fulton|east|rotor|75|shipped
1530|fulton|south|panel|20|shipped
1778|ember|east|valve|74|shipped
1793|ionic|south|sensor|59|paid
1606|acme|west|panel|95|held
1511|acme|north|frame|64|paid
1382|ember|north|frame|76|held
1301|cobalt|west|gasket|25|pending
1596|ionic|north|rotor|19|pending
1568|cobalt|west|frame|91|held
1663|acme|west|frame|95|held
1919|birch|west|rotor|66|paid
1296|cobalt|south|gasket|14|paid
1982|acme|east|cable|71|pending
1924|birch|south|panel|20|shipped
1335|acme|north|sensor|67|shipped
1285|cobalt|south|cable|43|paid
1923|gale|east|gasket|11|paid
1976|harbor|west|valve|51|paid
1457|dorian|south|pump|52|pending
1820|dorian|west|valve|19|held
2014|gale|north|frame|72|paid
1651|gale|west|pump|98|shipped
1579|juno|west|valve|23|paid
1498|ionic|south|sensor|98|paid
1866|gale|west|frame|68|paid
1987|fulton|west|frame|74|paid
1318|cobalt|south|frame|22|pending
1494|harbor|west|pump|21|shipped
1767|cobalt|west|sensor|68|held
1342|ionic|south|panel|58|held
1622|acme|south|sensor|53|shipped
1363|acme|north|sensor|57|pending
1833|fulton|west|frame|43|pending
1951|harbor|west|valve|61|held
1611|ember|east|frame|97|pending
1682|birch|west|rotor|64|held
1875|fulton|south|sensor|23|held
1483|cobalt|north|pump|61|paid
1370|dorian|west|frame|10|pending
1469|juno|south|valve|48|paid
1520|ionic|north|frame|42|shipped
1869|harbor|north|gasket|53|paid
1631|dorian|west|rotor|17|paid
1437|fulton|south|pump|93|shipped
2027|harbor|east|sensor|43|pending
1481|cobalt|south|valve|36|held
1855|birch|north|valve|83|paid
1763|harbor|north|valve|39|held
1953|gale|west|cable|90|pending
1497|harbor|west|panel|94|shipped
1376|ember|west|cable|74|shipped
1813|gale|north|rotor|47|paid
1535|harbor|west|rotor|86|paid
1699|gale|south|rotor|70|held
1523|dorian|south|panel|99|held
1447|harbor|east|frame|90|held
1397|dorian|west|valve|16|shipped
1769|juno|west|frame|70|held
1502|cobalt|west|sensor|69|pending
1897|cobalt|south|pump|74|pending
1541|cobalt|west|valve|66|pending
1747|harbor|east|sensor|47|held
1542|fulton|north|gasket|98|shipped
```
POLICY (apply exactly):
- Consider ONLY orders where customer = "cobalt" AND region = "south" AND status = "pending".
- Process those orders in ASCENDING order id.
- For each: if qty > 65, call restock(item, qty) with that order's item and qty; otherwise call cancel_order(order_id).
- Call nothing else. Every other record is irrelevant no matter how similar it looks.
Output the full ordered sequence of tool calls as a JSON array of {"tool": string, "args": object}.
Give your final answer inside ONE fenced code block (```), containing exactly the required content and nothing else. After the code block, add one plain-text line:
CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>model answer:
[
{"tool": "cancel_order", "args": {"order_id": 1271}},
{"tool": "cancel_order", "args": {"order_id": 1287}},
{"tool": "cancel_order", "args": {"order_id": 1297}},
{"tool": "restock", "args": {"item": "sensor", "qty": 71}},
{"tool": "cancel_order", "args": {"order_id": 1318}},
{"tool": "cancel_order", "args": {"order_id": 1598}},
{"tool": "restock", "args": {"item": "pump", "qty": 74}}
]correctagentic.tools.triage-v1conf 100% · 1.7s · $0.016 · 1744 tok
question
You operate a support desk strictly through tool calls.
TOOL CATALOG:
- create_ticket(title: string, priority: int) — opens a ticket; IDs are assigned sequentially: the 1st created ticket is "TCK-1", the 2nd "TCK-2", etc.
- assign(ticket_id: string, agent: string) — assigns an open ticket to an agent
- escalate(ticket_id: string) — marks a ticket as escalated
- close_ticket(ticket_id: string, resolution: string) — closes a ticket with a resolution note
- send_email(to: string, body: string) — sends an email (not part of the triage policy)
- archive_ticket(ticket_id: string) — archives a closed ticket (nightly job does this automatically)
- set_reminder(ticket_id: string, hours: int) — sets a follow-up reminder
ROUTING POLICY (apply exactly, in this order, for each incident, processing incidents in the order listed):
1. Create a ticket for the incident (title = the incident text verbatim, priority as given).
2. If the incident is an exact duplicate of an earlier incident in this list, close its ticket immediately with resolution "duplicate of <ID of the earlier ticket>" and do nothing else for it.
3. Otherwise, if priority ≥ 8, escalate the ticket BEFORE assigning it.
4. Assign the ticket to the agent responsible for its category.
CATEGORY → AGENT:
- infra → novak
- data → rivera
- auth → tanaka
INCIDENTS:
1. "webhooks not delivered" (category: infra, priority 2)
2. "webhooks not delivered" (category: infra, priority 2)
3. "cannot reset password" (category: auth, priority 6)
Output the full ordered sequence of tool calls needed to accomplish the goal, as a JSON array of objects {"tool": string, "args": object}. Use exactly the tool and argument names from the catalog. Do not call any tool that is not required.
Give your final answer inside ONE fenced code block (```), containing exactly the required content and nothing else. After the code block, add one plain-text line:
CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>model answer:
[
{
"tool": "create_ticket",
"args": {
"title": "webhooks not delivered",
"priority": 2
}
},
{
"tool": "assign",
"args": {
"ticket_id": "TCK-1",
"agent": "novak"
}
},
{
"tool": "create_ticket",
"args": {
"title": "webhooks not delivered",
"priority": 2
}
},
{
"tool": "close_ticket",
"args": {
"ticket_id": "TCK-2",
"resolution": "duplicate of TCK-1"
}
},
{
"tool": "create_ticket",
"args": {
"title": "cannot reset password",
"priority": 6
}
},
{
"tool": "assign",
"args": {
"ticket_id": "TCK-3",
"agent": "tanaka"
}
}
]correctagentic.tools.context-load-v1anchorconf 100% · 1.4s · $0.089 · 9233 tok
model answer:
[
{"tool": "restock", "args": {"item": "gasket", "qty": 60}},
{"tool": "cancel_order", "args": {"order_id": 1368}},
{"tool": "cancel_order", "args": {"order_id": 1384}},
{"tool": "restock", "args": {"item": "frame", "qty": 58}},
{"tool": "restock", "args": {"item": "rotor", "qty": 75}},
{"tool": "cancel_order", "args": {"order_id": 1682}}
]correctagentic.tools.deploy-v1conf 100% · 1.7s · $0.011 · 1158 tok
question
You are a release agent operating strictly through tool calls.
TOOL CATALOG:
- deploy(service: string) — deploys a service; FAILS if any dependency is not yet deployed
- health_check(service: string) — runs a post-deploy health probe
- rollback(service: string) — reverts a bad deploy (nothing fails in this scenario)
- scale(service: string, replicas: int) — changes replica count (out of scope)
- restart(service: string) — restarts a service (out of scope)
SERVICES AND THEIR DEPENDENCIES (a service can only be deployed after ALL its dependencies):
- gateway: reports
- notifier: reports
- search: gateway
- reports: (none)
POLICY:
- Deploy in waves: in each wave, deploy every service whose dependencies are already deployed, in alphabetical order; repeat until all services are deployed.
- The service "search" is flagged unstable: call health_check on it immediately after deploying it.
- Call nothing else.
Output the full ordered sequence of tool calls needed to accomplish the goal, as a JSON array of objects {"tool": string, "args": object}. Use exactly the tool and argument names from the catalog. Do not call any tool that is not required.
Give your final answer inside ONE fenced code block (```), containing exactly the required content and nothing else. After the code block, add one plain-text line:
CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>model answer:
[
{
"tool": "deploy",
"args": {
"service": "reports"
}
},
{
"tool": "deploy",
"args": {
"service": "gateway"
}
},
{
"tool": "deploy",
"args": {
"service": "notifier"
}
},
{
"tool": "deploy",
"args": {
"service": "search"
}
},
{
"tool": "health_check",
"args": {
"service": "search"
}
}
]correctagentic.tools.ledger-v1conf 100% · 1.4s · $0.018 · 1881 tok
question
You are a treasury agent operating strictly through tool calls.
TOOL CATALOG:
- transfer(from: string, to: string, amount: int) — moves funds between accounts; FAILS if it would overdraw the source
- top_up_from_reserve(account: string, amount: int) — adds funds to an account from the corporate reserve
- get_balance(account: string) — reads a balance (you already have all balances below — reads are unnecessary and forbidden)
- freeze_account(account: string) — compliance freeze (not part of this task)
- convert_currency(account: string, currency: string) — FX conversion (all amounts are already in USD)
OPENING BALANCES:
- tango: $760
- echo: $896
- kilo: $312
PAYMENT INSTRUCTIONS (execute in exactly this order):
1. pay $93 from "kilo" to "tango"
2. pay $512 from "tango" to "echo"
3. pay $254 from "tango" to "kilo"
4. pay $267 from "kilo" to "tango"
RULES:
- transfer() fails on overdraft. If a payment would overdraw its source account at the moment of execution, first call top_up_from_reserve() on the source with EXACTLY the shortfall (no more, no less), then execute the transfer.
- Track balances as they change: earlier payments affect later ones.
- Never call tools that are not needed.
Output the full ordered sequence of tool calls needed to accomplish the goal, as a JSON array of objects {"tool": string, "args": object}. Use exactly the tool and argument names from the catalog. Do not call any tool that is not required.
Give your final answer inside ONE fenced code block (```), containing exactly the required content and nothing else. After the code block, add one plain-text line:
CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>model answer:
[
{
"tool": "transfer",
"args": {
"from": "kilo",
"to": "tango",
"amount": 93
}
},
{
"tool": "transfer",
"args": {
"from": "tango",
"to": "echo",
"amount": 512
}
},
{
"tool": "transfer",
"args": {
"from": "tango",
"to": "kilo",
"amount": 254
}
},
{
"tool": "transfer",
"args": {
"from": "kilo",
"to": "tango",
"amount": 267
}
}
]wrongagentic.tools.triage-v1conf — · 2.2s · $0.000 · 0 tok
question
You operate a support desk strictly through tool calls.
TOOL CATALOG:
- create_ticket(title: string, priority: int) — opens a ticket; IDs are assigned sequentially: the 1st created ticket is "TCK-1", the 2nd "TCK-2", etc.
- assign(ticket_id: string, agent: string) — assigns an open ticket to an agent
- escalate(ticket_id: string) — marks a ticket as escalated
- close_ticket(ticket_id: string, resolution: string) — closes a ticket with a resolution note
- send_email(to: string, body: string) — sends an email (not part of the triage policy)
- archive_ticket(ticket_id: string) — archives a closed ticket (nightly job does this automatically)
- set_reminder(ticket_id: string, hours: int) — sets a follow-up reminder
ROUTING POLICY (apply exactly, in this order, for each incident, processing incidents in the order listed):
1. Create a ticket for the incident (title = the incident text verbatim, priority as given).
2. If the incident is an exact duplicate of an earlier incident in this list, close its ticket immediately with resolution "duplicate of <ID of the earlier ticket>" and do nothing else for it.
3. Otherwise, if priority ≥ 7, escalate the ticket BEFORE assigning it.
4. Assign the ticket to the agent responsible for its category.
CATEGORY → AGENT:
- infra → okafor
- auth → haddad
- payments → dubois
INCIDENTS:
1. "webhooks not delivered" (category: infra, priority 5)
2. "webhooks not delivered" (category: infra, priority 5)
3. "invoice total wrong" (category: payments, priority 8)
Output the full ordered sequence of tool calls needed to accomplish the goal, as a JSON array of objects {"tool": string, "args": object}. Use exactly the tool and argument names from the catalog. Do not call any tool that is not required.
Give your final answer inside ONE fenced code block (```), containing exactly the required content and nothing else. After the code block, add one plain-text line:
CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>model answer:
(none extracted)correctagentic.tools.deploy-v1anchorconf 100% · 8.9s · $0.012 · 1272 tok
model answer:
[
{
"tool": "deploy",
"args": {
"service": "notifier"
}
},
{
"tool": "deploy",
"args": {
"service": "auth-svc"
}
},
{
"tool": "deploy",
"args": {
"service": "gateway"
}
},
{
"tool": "health_check",
"args": {
"service": "gateway"
}
},
{
"tool": "deploy",
"args": {
"service": "reports"
}
}
]correctagentic.tools.triage-v1anchorconf 100% · 1.4s · $0.021 · 2275 tok
model answer:
[
{
"tool": "create_ticket",
"args": {
"title": "API latency spikes",
"priority": 9
}
},
{
"tool": "escalate",
"args": {
"ticket_id": "TCK-1"
}
},
{
"tool": "assign",
"args": {
"ticket_id": "TCK-1",
"agent": "tanaka"
}
},
{
"tool": "create_ticket",
"args": {
"title": "records missing after import",
"priority": 2
}
},
{
"tool": "assign",
"args": {
"ticket_id": "TCK-2",
"agent": "novak"
}
},
{
"tool": "create_ticket",
"args": {
"title": "refund double-charged",
"priority": 2
}
},
{
"tool": "assign",
"args": {
"ticket_id": "TCK-3",
"agent": "okafor"
}
},
{
"tool": "create_ticket",
"args": {
"title": "API latency spikes",
"priority": 9
}
},
{
"tool": "close_ticket",
"args": {
"ticket_id": "TCK-4",
"resolution": "duplicate of TCK-1"
}
}
]correctagentic.tools.ledger-v1anchorconf 100% · 1.5s · $0.019 · 2004 tok
model answer:
[
{"tool": "top_up_from_reserve", "args": {"account": "delta", "amount": 134}},
{"tool": "transfer", "args": {"from": "delta", "to": "bravo", "amount": 451}},
{"tool": "transfer", "args": {"from": "bravo", "to": "delta", "amount": 292}},
{"tool": "transfer", "args": {"from": "delta", "to": "tango", "amount": 97}},
{"tool": "top_up_from_reserve", "args": {"account": "delta", "amount": 345}},
{"tool": "transfer", "args": {"from": "delta", "to": "bravo", "amount": 540}},
{"tool": "transfer", "args": {"from": "tango", "to": "bravo", "amount": 125}}
]code 30/30 correct
correctcode.trace.js-v1conf 100% · 10.0s · $0.013 · 1395 tok
question
Evaluate the following JavaScript. What number is logged to the console? ```js const arr = [6, 7, 8, 9, 10, 11, 12, 13, 14, 15]; const out = arr .map(n => n * 6) .filter(n => n % 4 === 0) .reduce((a, b) => a + b, 0); console.log(out); ``` End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
300correctcode.trace.nested-v1conf 100% · 1.6s · $0.034 · 3755 tok
question
Trace this Python program exactly. What does it print?
```python
total = 0
for i in range(1, 8):
for j in range(1, 6):
if j == 3 and i % 2 == 0:
break
if (i + j) % 3 == 0:
continue
total += i * 3 + j
print(total)
```
End your reply with exactly two plain-text lines (no markdown, no extra text after them):
ANSWER: <your final answer only>
CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>model answer:
251correctcode.trace.python-v1conf 100% · 1.7s · $0.015 · 1677 tok
question
Execute this Python snippet mentally. What is printed?
```python
total = 0
v = 12
while total + v <= 109:
if v % 3 != 0:
total += v
v += 5
print(total)
```
End your reply with exactly two plain-text lines (no markdown, no extra text after them):
ANSWER: <your final answer only>
CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>model answer:
108correctcode.trace.nested-v1conf 100% · 10.0s · $0.029 · 3197 tok
question
Trace this Python program exactly. What does it print?
```python
total = 0
for i in range(1, 5):
for j in range(1, 7):
if j == 5 and i % 2 == 0:
break
if (i + j) % 4 == 0:
continue
total += i * 3 + j
print(total)
```
End your reply with exactly two plain-text lines (no markdown, no extra text after them):
ANSWER: <your final answer only>
CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>model answer:
152correctcode.trace.js-v1conf 100% · 1.8s · $0.010 · 1067 tok
question
What does this JavaScript program log? ```js const arr = [5, 6, 7, 8, 9, 10, 11, 12, 13, 14]; const out = arr .map(n => n * 3) .filter(n => n % 5 === 0) .reduce((a, b) => a + b, 0); console.log(out); ``` End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
45correctcode.trace.python-v1conf 100% · 1.7s · $0.027 · 2925 tok
question
Trace the following Python code and give its exact output.
```python
total = 0
v = 2
while total + v <= 106:
if v % 6 != 0:
total += v
v += 4
print(total)
```
End your reply with exactly two plain-text lines (no markdown, no extra text after them):
ANSWER: <your final answer only>
CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>model answer:
74correctcode.trace.nested-v1conf 100% · 1.6s · $0.023 · 2476 tok
question
Trace this Python program exactly. What does it print?
```python
total = 0
for i in range(1, 7):
for j in range(1, 5):
if j == 3 and i % 2 == 0:
break
if (i + j) % 2 == 0:
continue
total += i * 3 + j
print(total)
```
End your reply with exactly two plain-text lines (no markdown, no extra text after them):
ANSWER: <your final answer only>
CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>model answer:
111correctcode.trace.js-v1conf 100% · 1.9s · $0.008 · 897 tok
question
Evaluate the following JavaScript. What number is logged to the console? ```js const arr = [5, 6, 7, 8, 9, 10, 11, 12]; const out = arr .map(n => n * 2) .filter(n => n % 2 === 0) .reduce((a, b) => a + b, 0); console.log(out); ``` End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
136correctcode.trace.python-v1conf 100% · 1.6s · $0.014 · 1588 tok
question
Execute this Python snippet mentally. What is printed?
```python
total = 0
v = 6
while total + v <= 96:
if v % 6 != 0:
total += v
v += 2
print(total)
```
End your reply with exactly two plain-text lines (no markdown, no extra text after them):
ANSWER: <your final answer only>
CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>model answer:
90correctcode.trace.nested-v1conf 100% · 7.6s · $0.030 · 3318 tok
question
Trace this Python program exactly. What does it print?
```python
total = 0
for i in range(1, 5):
for j in range(1, 7):
if j == 3 and i % 2 == 0:
break
if (i + j) % 2 == 0:
continue
total += i * 4 + j
print(total)
```
End your reply with exactly two plain-text lines (no markdown, no extra text after them):
ANSWER: <your final answer only>
CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>model answer:
98correctcode.trace.js-v1conf 100% · 1.7s · $0.009 · 936 tok
question
What does this JavaScript program log? ```js const arr = [3, 4, 5, 6, 7, 8, 9, 10, 11, 12]; const out = arr .map(n => n * 7) .filter(n => n % 4 === 0) .reduce((a, b) => a + b, 0); console.log(out); ``` End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
168correctcode.trace.python-v1conf 100% · 7.7s · $0.009 · 945 tok
question
What does this Python program print?
```python
total = 0
v = 4
while total + v <= 38:
if v % 3 != 0:
total += v
v += 7
print(total)
```
End your reply with exactly two plain-text lines (no markdown, no extra text after them):
ANSWER: <your final answer only>
CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>model answer:
15correctcode.trace.nested-v1conf 100% · 1.4s · $0.030 · 3273 tok
question
Trace this Python program exactly. What does it print?
```python
total = 0
for i in range(1, 7):
for j in range(1, 5):
if j == 3 and i % 2 == 0:
break
if (i + j) % 3 == 0:
continue
total += i * 3 + j
print(total)
```
End your reply with exactly two plain-text lines (no markdown, no extra text after them):
ANSWER: <your final answer only>
CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>model answer:
146correctcode.trace.js-v1conf 100% · 3.1s · $0.004 · 468 tok
question
What does this JavaScript program log? ```js const arr = [1, 2, 3, 4, 5, 6]; const out = arr .map(n => n * 2) .filter(n => n % 3 === 0) .reduce((a, b) => a + b, 0); console.log(out); ``` End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
18correctcode.trace.python-v1conf 100% · 1.5s · $0.014 · 1510 tok
question
What does this Python program print?
```python
total = 0
v = 3
while total + v <= 47:
if v % 6 != 0:
total += v
v += 4
print(total)
```
End your reply with exactly two plain-text lines (no markdown, no extra text after them):
ANSWER: <your final answer only>
CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>model answer:
36correctcode.trace.nested-v1conf 100% · 1.4s · $0.021 · 2279 tok
question
Trace this Python program exactly. What does it print?
```python
total = 0
for i in range(1, 5):
for j in range(1, 6):
if j == 4 and i % 2 == 0:
break
if (i + j) % 2 == 0:
continue
total += i * 5 + j
print(total)
```
End your reply with exactly two plain-text lines (no markdown, no extra text after them):
ANSWER: <your final answer only>
CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>model answer:
120correctcode.trace.js-v1conf 100% · 1.6s · $0.006 · 620 tok
question
What does this JavaScript program log? ```js const arr = [5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16]; const out = arr .map(n => n * 3) .filter(n => n % 5 === 0) .reduce((a, b) => a + b, 0); console.log(out); ``` End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
90correctcode.trace.python-v1conf 100% · 1.5s · $0.013 · 1435 tok
question
Execute this Python snippet mentally. What is printed?
```python
total = 0
v = 14
while total + v <= 90:
if v % 3 != 0:
total += v
v += 6
print(total)
```
End your reply with exactly two plain-text lines (no markdown, no extra text after them):
ANSWER: <your final answer only>
CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>model answer:
60correctcode.trace.nested-v1conf 100% · 1.5s · $0.032 · 3554 tok
question
Trace this Python program exactly. What does it print?
```python
total = 0
for i in range(1, 8):
for j in range(1, 7):
if j == 3 and i % 2 == 0:
break
if (i + j) % 2 == 0:
continue
total += i * 3 + j
print(total)
```
End your reply with exactly two plain-text lines (no markdown, no extra text after them):
ANSWER: <your final answer only>
CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>model answer:
231correctcode.trace.js-v1conf 100% · 1.6s · $0.010 · 1128 tok
question
Evaluate the following JavaScript. What number is logged to the console? ```js const arr = [4, 5, 6, 7, 8, 9, 10, 11]; const out = arr .map(n => n * 3) .filter(n => n % 4 === 0) .reduce((a, b) => a + b, 0); console.log(out); ``` End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
36correctcode.trace.python-v1conf 100% · 1.7s · $0.021 · 2294 tok
question
What does this Python program print?
```python
total = 0
v = 14
while total + v <= 106:
if v % 7 != 0:
total += v
v += 4
print(total)
```
End your reply with exactly two plain-text lines (no markdown, no extra text after them):
ANSWER: <your final answer only>
CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>model answer:
96correctcode.trace.js-v1conf 100% · 10.1s · $0.010 · 1104 tok
question
What does this JavaScript program log? ```js const arr = [5, 6, 7, 8, 9, 10, 11, 12, 13, 14]; const out = arr .map(n => n * 6) .filter(n => n % 2 === 0) .reduce((a, b) => a + b, 0); console.log(out); ``` End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
570correctcode.trace.nested-v1conf 100% · 1.7s · $0.029 · 3244 tok
question
Trace this Python program exactly. What does it print?
```python
total = 0
for i in range(1, 6):
for j in range(1, 7):
if j == 5 and i % 2 == 0:
break
if (i + j) % 2 == 0:
continue
total += i * 3 + j
print(total)
```
End your reply with exactly two plain-text lines (no markdown, no extra text after them):
ANSWER: <your final answer only>
CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>model answer:
161correctcode.trace.nested-v1conf 100% · 8.0s · $0.027 · 2953 tok
question
Trace this Python program exactly. What does it print?
```python
total = 0
for i in range(1, 5):
for j in range(1, 6):
if j == 3 and i % 2 == 0:
break
if (i + j) % 3 == 0:
continue
total += i * 4 + j
print(total)
```
End your reply with exactly two plain-text lines (no markdown, no extra text after them):
ANSWER: <your final answer only>
CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>model answer:
107correctcode.trace.python-v1conf 100% · 1.9s · $0.009 · 938 tok
question
Execute this Python snippet mentally. What is printed?
```python
total = 0
v = 3
while total + v <= 39:
if v % 3 != 0:
total += v
v += 8
print(total)
```
End your reply with exactly two plain-text lines (no markdown, no extra text after them):
ANSWER: <your final answer only>
CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>model answer:
30correctcode.trace.js-v1conf 100% · 2.1s · $0.010 · 1115 tok
question
Evaluate the following JavaScript. What number is logged to the console? ```js const arr = [7, 8, 9, 10, 11, 12, 13, 14, 15, 16]; const out = arr .map(n => n * 5) .filter(n => n % 3 === 0) .reduce((a, b) => a + b, 0); console.log(out); ``` End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
180correctcode.trace.python-v1anchorconf 100% · 1.7s · $0.012 · 1336 tok
model answer:
0correctcode.trace.nested-v1anchorconf 100% · 8.2s · $0.030 · 3264 tok
model answer:
283correctcode.trace.js-v1anchorconf 100% · 8.6s · $0.009 · 968 tok
model answer:
63correctcode.trace.python-v1anchorconf 100% · 2.5s · $0.009 · 963 tok
model answer:
40instruction following 30/30 correct
correctif.format.repeat-v1conf 100% · 2.7s · $0.004 · 376 tok
question
Write the word "falcon" in uppercase form, repeated exactly 3 times, joined by the character "-" with no spaces. Output that string as your answer. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
FALCON-FALCON-FALCONcorrectif.constraints.stack-v1conf 100% · 1.5s · $0.021 · 2275 tok
question
Write in English about an old machine, following ALL of these rules simultaneously: 1. Exactly 15 words. 2. The first word must be "ember" and the last word must be "lumen". 3. Use the word "quartz" exactly 2 times (in addition to rules 2 if they differ). 4. The letter "j" must not appear anywhere. 5. Everything entirely in lowercase. Give the text itself as your answer. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
ember sparks this old quartz engine as each quartz gear clicks to release its lumencorrectif.constraints.stack-v1conf 100% · 10.0s · $0.026 · 2887 tok
question
Write in English about an old machine, following ALL of these rules simultaneously: 1. Exactly 20 words. 2. The first word must be "comet" and the last word must be "falcon". 3. Use the word "quartz" exactly 2 times (in addition to rules 2 if they differ). 4. The letter "j" must not appear anywhere. 5. Everything entirely in lowercase. Give the text itself as your answer. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
comet was an old quartz engine that spun gears inside the dusty mill near quartz mines under a metal falconcorrectif.format.acronym-v1conf 100% · 2.2s · $0.003 · 350 tok
question
Take the first letter of each of these words, in order: delta, comet, ember, orbit, zephyr, drift. Concatenate them in uppercase into a single string with no separators. Output that string as your answer. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
DCEOZDcorrectif.constraints.stack-v1conf 100% · 1.5s · $0.022 · 2453 tok
question
Write in English about a long journey, following ALL of these rules simultaneously: 1. Exactly 19 words. 2. The first word must be "basalt" and the last word must be "echo". 3. Use the word "prism" exactly 2 times (in addition to rules 2 if they differ). 4. The letter "q" must not appear anywhere. 5. Everything entirely in lowercase. Give the text itself as your answer. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
basalt cliffs faded as we carried one prism across endless deserts to find another prism and its lost echocorrectif.format.acronym-v1conf 100% · 1.6s · $0.003 · 347 tok
question
Take the first letter of each of these words, in order: basalt, tundra, quartz, prism, cedar. Concatenate them in uppercase into a single string with no separators. Output that string as your answer. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
BTQPCcorrectif.format.repeat-v1conf 100% · 1.4s · $0.010 · 1100 tok
question
Write the word "cedar" in uppercase form, repeated exactly 9 times, joined by the character "-" with no spaces. Output that string as your answer. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
CEDAR-CEDAR-CEDAR-CEDAR-CEDAR-CEDAR-CEDAR-CEDAR-CEDARcorrectif.format.repeat-v1conf 95% · 1.9s · $0.023 · 2549 tok
question
Write the word "zephyr" in capitalized form, repeated exactly 7 times, joined by the character "-" with no spaces. Output that string as your answer. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Zephyr-Zephyr-Zephyr-Zephyr-Zephyr-Zephyr-Zephyrcorrectif.constraints.stack-v1conf 100% · 1.8s · $0.027 · 2972 tok
question
Write in English about an old machine, following ALL of these rules simultaneously: 1. Exactly 16 words. 2. The first word must be "lumen" and the last word must be "cedar". 3. Use the word "comet" exactly 3 times (in addition to rules 2 if they differ). 4. The letter "q" must not appear anywhere. 5. Everything entirely in lowercase. Give the text itself as your answer. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
lumen by lumen the old comet engine spun a comet toward a comet built of cedarcorrectif.format.acronym-v1conf 100% · 1.5s · $0.004 · 383 tok
question
Take the first letter of each of these words, in order: zephyr, flint, tundra, falcon, comet, ember. Concatenate them in uppercase into a single string with no separators. Output that string as your answer. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
ZFTFCEcorrectif.format.acronym-v1conf 100% · 1.6s · $0.003 · 362 tok
question
Take the first letter of each of these words, in order: echo, falcon, ember, cedar, flint. Concatenate them in uppercase into a single string with no separators. Output that string as your answer. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
EFECFcorrectif.format.repeat-v1conf 90% · 2.8s · $0.039 · 4269 tok
question
Write the word "comet" in capitalized form, repeated exactly 5 times, joined by the character "_" with no spaces. Output that string as your answer. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
COMET_COMET_COMET_COMET_COMETcorrectif.constraints.stack-v1conf 100% · 1.6s · $0.024 · 2629 tok
question
Write in English about a city at night, following ALL of these rules simultaneously: 1. Exactly 14 words. 2. The first word must be "delta" and the last word must be "tundra". 3. Use the word "basalt" exactly 3 times (in addition to rules 2 if they differ). 4. The letter "q" must not appear anywhere. 5. Everything entirely in lowercase. Give the text itself as your answer. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
delta lights shine over basalt walls as basalt roads reflect basalt under dark tundracorrectif.format.repeat-v1conf 100% · 1.6s · $0.011 · 1246 tok
question
Write the word "basalt" in lowercase form, repeated exactly 9 times, joined by the character "/" with no spaces. Output that string as your answer. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
basalt/basalt/basalt/basalt/basalt/basalt/basalt/basalt/basaltcorrectif.format.acronym-v1conf 100% · 8.0s · $0.005 · 552 tok
question
Take the third letter of each of these words, in order: echo, orbit, basalt, zephyr, lumen, ember. Concatenate them in uppercase into a single string with no separators. Output that string as your answer. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
HBSPMBcorrectif.constraints.stack-v1conf 100% · 2.0s · $0.022 · 2401 tok
question
Write in English about the sea, following ALL of these rules simultaneously: 1. Exactly 20 words. 2. The first word must be "ember" and the last word must be "drift". 3. Use the word "echo" exactly 2 times (in addition to rules 2 if they differ). 4. The letter "q" must not appear anywhere. 5. Everything entirely in lowercase. Give the text itself as your answer. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
ember of the sun sinks low where waves echo and tides echo through deep blue waters as silent ships driftcorrectif.format.acronym-v1conf 100% · 7.8s · $0.004 · 377 tok
question
Take the second letter of each of these words, in order: delta, falcon, basalt, tundra. Concatenate them in uppercase into a single string with no separators. Output that string as your answer. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
EAAUcorrectif.constraints.stack-v1conf 100% · 1.5s · $0.038 · 4156 tok
question
Write in English about a long journey, following ALL of these rules simultaneously: 1. Exactly 19 words. 2. The first word must be "prism" and the last word must be "orbit". 3. Use the word "cedar" exactly 3 times (in addition to rules 2 if they differ). 4. The letter "q" must not appear anywhere. 5. Everything entirely in lowercase. Give the text itself as your answer. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
prism of light guided our long journey past cedar trees where cedar logs and cedar smoke rose toward orbitcorrectif.format.repeat-v1conf 100% · 1.5s · $0.008 · 843 tok
question
Write the word "prism" in lowercase form, repeated exactly 7 times, joined by the character "-" with no spaces. Output that string as your answer. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
prism-prism-prism-prism-prism-prism-prismcorrectif.constraints.stack-v1conf 100% · 1.5s · $0.021 · 2289 tok
question
Write in English about the sea, following ALL of these rules simultaneously: 1. Exactly 16 words. 2. The first word must be "delta" and the last word must be "falcon". 3. Use the word "flint" exactly 2 times (in addition to rules 2 if they differ). 4. The letter "q" must not appear anywhere. 5. Everything entirely in lowercase. Give the text itself as your answer. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
delta waves crash on the shore where grey flint meets cold flint under a flying falconcorrectif.format.acronym-v1conf 100% · 1.5s · $0.004 · 414 tok
question
Take the first letter of each of these words, in order: ember, basalt, falcon, echo, lumen, quartz. Concatenate them in uppercase into a single string with no separators. Output that string as your answer. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
EBFELQcorrectif.format.repeat-v1conf 100% · 1.7s · $0.006 · 617 tok
question
Write the word "zephyr" in uppercase form, repeated exactly 3 times, joined by the character "-" with no spaces. Output that string as your answer. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
ZEPHYR-ZEPHYR-ZEPHYRcorrectif.format.acronym-v1conf 100% · 1.7s · $0.005 · 542 tok
question
Take the third letter of each of these words, in order: nova, quartz, lumen, echo, cedar, zephyr. Concatenate them in uppercase into a single string with no separators. Output that string as your answer. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
VAMHDPcorrectif.constraints.stack-v1conf 100% · 1.6s · $0.019 · 2112 tok
question
Write in English about a long journey, following ALL of these rules simultaneously: 1. Exactly 23 words. 2. The first word must be "falcon" and the last word must be "delta". 3. Use the word "echo" exactly 3 times (in addition to rules 2 if they differ). 4. The letter "q" must not appear anywhere. 5. Everything entirely in lowercase. Give the text itself as your answer. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
falcon flew over mountains hearing an echo then another echo until a third echo guided his long flight to the warm river deltacorrectif.format.repeat-v1conf 100% · 1.4s · $0.009 · 1036 tok
question
Write the word "prism" in lowercase form, repeated exactly 9 times, joined by the character "_" with no spaces. Output that string as your answer. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
prism_prism_prism_prism_prism_prism_prism_prism_prismcorrectif.format.repeat-v1anchorconf 100% · 10.0s · $0.028 · 3133 tok
model answer:
Echo_Echo_Echo_Echo_Echo_Echo_Echo_Echocorrectif.constraints.stack-v1anchorconf 100% · 10.0s · $0.037 · 4062 tok
model answer:
drift through the dark neon streets where lumen meets lumen under a soft lumen of starlight as silent towers watch their slow eternal orbitcorrectif.format.acronym-v1conf 100% · 1.5s · $0.005 · 537 tok
question
Take the third letter of each of these words, in order: orbit, zephyr, lumen, comet. Concatenate them in uppercase into a single string with no separators. Output that string as your answer. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
BPMMcorrectif.format.repeat-v1anchorconf 100% · 1.6s · $0.008 · 868 tok
model answer:
ZEPHYR/ZEPHYR/ZEPHYR/ZEPHYR/ZEPHYR/ZEPHYRcorrectif.format.acronym-v1anchorconf 100% · 1.6s · $0.004 · 402 tok
model answer:
ZDFQknowledge 30/30 correct
correctknowledge.fr.factbank-v2conf 100% · 1.5s · $0.004 · 457 tok
question
Name the element whose symbol is Hg. Answer with the name only — no explanation. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Mercurycorrectknowledge.fr.factbank-v2conf 100% · 7.6s · $0.006 · 677 tok
question
Name the chemical element with symbol Hg. Answer with the name only — no explanation. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Mercurycorrectknowledge.fr.factbank-v2conf 100% · 10.0s · $0.003 · 343 tok
question
Identify the element whose symbol is Hg. Answer with the name only — no explanation. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Mercurycorrectknowledge.fr.factbank-v2conf 100% · 9.3s · $0.008 · 901 tok
question
Identify the chemical element with symbol Sn. Answer with the name only — no explanation. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Tincorrectknowledge.fr.factbank-v2conf 100% · 1.6s · $0.007 · 728 tok
question
Identify the Swiss capital (de facto). Answer with the name only — no explanation. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Berncorrectknowledge.fr.factbank-v2conf 100% · 1.7s · $0.006 · 600 tok
question
What is the writer of the novel "One Hundred Years of Solitude"? Answer with the name only — no explanation. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Gabriel García Márquezcorrectknowledge.fr.factbank-v2conf 100% · 1.8s · $0.002 · 164 tok
question
Identify the chemical element with symbol Sb. Answer with the name only — no explanation. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Antimonycorrectknowledge.fr.factbank-v2conf 100% · 2.1s · $0.007 · 756 tok
question
Identify the Burmese capital city. Answer with the name only — no explanation. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Naypyidawcorrectknowledge.fr.factbank-v2conf 100% · 1.9s · $0.007 · 806 tok
question
Name the capital of Switzerland. Answer with the name only — no explanation. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Berncorrectknowledge.fr.factbank-v2conf 100% · 1.4s · $0.009 · 1038 tok
question
Name the author of "The Master and Margarita". Answer with the name only — no explanation. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Mikhail Bulgakovcorrectknowledge.fr.factbank-v2conf 100% · 1.5s · $0.006 · 653 tok
question
Name the element whose symbol is Hg. Answer with the name only — no explanation. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Mercurycorrectknowledge.fr.factbank-v2conf 100% · 7.8s · $0.006 · 686 tok
question
Name the element whose symbol is K. Answer with the name only — no explanation. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Potassiumcorrectknowledge.fr.factbank-v2conf 100% · 1.5s · $0.002 · 178 tok
question
What is the chemical element with symbol Sn? Answer with the name only — no explanation. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Tincorrectknowledge.fr.factbank-v2conf 100% · 1.2s · $0.007 · 712 tok
question
Identify the Brazilian capital city. Answer with the name only — no explanation. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Brasíliacorrectknowledge.fr.factbank-v2conf 100% · 2.1s · $0.004 · 451 tok
question
Name the element whose symbol is Hg. Answer with the name only — no explanation. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Mercurycorrectknowledge.fr.factbank-v2conf 100% · 1.4s · $0.006 · 613 tok
question
Identify the Turkish capital city. Answer with the name only — no explanation. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Ankaracorrectknowledge.fr.factbank-v2conf 100% · 1.5s · $0.007 · 741 tok
question
Identify the Nigerian capital city. Answer with the name only — no explanation. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Abujacorrectknowledge.fr.factbank-v2conf 100% · 1.6s · $0.008 · 868 tok
question
What is the Brazilian capital city? Answer with the name only — no explanation. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Brasíliacorrectknowledge.fr.factbank-v2conf 100% · 1.5s · $0.004 · 464 tok
question
Name the Turkish capital city. Answer with the name only — no explanation. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Ankaracorrectknowledge.fr.factbank-v2conf 100% · 1.6s · $0.006 · 629 tok
question
Name the Kazakh capital city. Answer with the name only — no explanation. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Astanacorrectknowledge.fr.factbank-v2conf 100% · 1.5s · $0.004 · 390 tok
question
What is the capital of Brazil? Answer with the name only — no explanation. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Brasíliacorrectknowledge.fr.factbank-v2conf 100% · 1.6s · $0.008 · 867 tok
question
Name the writer of the novel "One Hundred Years of Solitude". Answer with the name only — no explanation. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Gabriel García Márquezcorrectknowledge.fr.factbank-v2conf 100% · 1.6s · $0.010 · 1107 tok
question
What is the capital of Myanmar? Answer with the name only — no explanation. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Naypyidawcorrectknowledge.fr.factbank-v2conf 100% · 10.0s · $0.007 · 798 tok
question
Identify the capital of Australia. Answer with the name only — no explanation. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Canberracorrectknowledge.fr.factbank-v2conf 100% · 1.7s · $0.006 · 667 tok
question
What is the chemical element with symbol Pb? Answer with the name only — no explanation. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Leadcorrectknowledge.fr.factbank-v2conf 100% · 1.8s · $0.007 · 713 tok
question
Identify the Nigerian capital city. Answer with the name only — no explanation. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Abujacorrectknowledge.fr.factbank-v2anchorconf 100% · 1.9s · $0.005 · 596 tok
model answer:
Mercurycorrectknowledge.fr.factbank-v2anchorconf 100% · 1.7s · $0.005 · 528 tok
model answer:
Tungstencorrectknowledge.fr.factbank-v2anchorconf 100% · 2.7s · $0.009 · 1029 tok
model answer:
Antimonycorrectknowledge.fr.factbank-v2anchorconf 100% · 7.9s · $0.001 · 144 tok
model answer:
Leadmath 30/30 correct
correctmath.chained.pipeline-v1conf 100% · 10.0s · $0.006 · 623 tok
question
Solve the following linked steps; each step uses the previous result. Step 1: P = 32 × 25. Step 2: Q = P × 5 − 317. Step 3: divide Q by 7: let q be the integer quotient and r the remainder. The final answer is q + r. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
527correctmath.counterfactual.base-v1conf 100% · 10.0s · $0.013 · 1462 tok
question
Work strictly in base 9. Multiply the base-9 numbers 85 and 107. Give the result IN BASE 9. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
10258correctmath.percent.chain-v2conf 100% · 1.6s · $0.010 · 1053 tok
question
An inventory starts at 10000 units. A rival firm shipped 61 unrelated parcels the same week. In the first month the inventory grows by 14%. Each pallet weighs about 46 grams more when wet. The next month it shrinks by 12%, and the month after it grows by 36%. How many units remain (exact value, round to 2 decimals only if needed)? End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
13643.52correctmath.algebra.system-v2conf 100% · 1.7s · $0.004 · 435 tok
question
Solve the system, then answer the derived question. 2x + 3y = 166 5x − 5y = 40 What is the value of 4x − 5y? End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
2correctmath.arith.chain-v2conf 100% · 1.6s · $0.011 · 1153 tok
question
Compute the value of the following expression. (((34 × 89 − 737) × 9 + 7936) − 75 × 81) × 7 End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
157234correctmath.chained.pipeline-v1conf 100% · 1.7s · $0.007 · 754 tok
question
Solve the following linked steps; each step uses the previous result. Step 1: P = 89 × 60. Step 2: Q = P × 8 − 676. Step 3: divide Q by 3: let q be the integer quotient and r the remainder. The final answer is q + r. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
14016correctmath.counterfactual.base-v1conf 100% · 1.9s · $0.011 · 1263 tok
question
Work strictly in base 8. Add the base-8 numbers 1074 and 337. Give the result IN BASE 8. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
1433correctmath.percent.chain-v2conf 100% · 1.3s · $0.011 · 1154 tok
question
An inventory starts at 16000 units. A rival firm shipped 19 unrelated parcels the same week. In the first month the inventory grows by 35%. A rival firm shipped 13 unrelated parcels the same week. The next month it shrinks by 31%, and the month after it grows by 26%. How many units remain (exact value, round to 2 decimals only if needed)? End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
18779.04correctmath.arith.chain-v2conf 100% · 1.5s · $0.011 · 1219 tok
question
Evaluate the expression below and give the result. (((66 × 26 − 313) × 3 + 1102) − 88 × 53) × 5 End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
3235correctmath.algebra.system-v2conf 100% · 2.5s · $0.004 · 406 tok
question
Solve the system, then answer the derived question. 3x + 5y = -79 8x − 5y = 101 What is the value of 2x − 4y? End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
72correctmath.chained.pipeline-v1conf 100% · 1.9s · $0.006 · 645 tok
question
Solve the following linked steps; each step uses the previous result. Step 1: P = 63 × 13. Step 2: Q = P × 3 − 540. Step 3: divide Q by 8: let q be the integer quotient and r the remainder. The final answer is q + r. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
244correctmath.counterfactual.base-v1conf 100% · 1.5s · $0.011 · 1253 tok
question
Work strictly in base 8. Add the base-8 numbers 2654 and 2166. Give the result IN BASE 8. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
5042correctmath.percent.chain-v2conf 100% · 1.5s · $0.013 · 1428 tok
question
An inventory starts at 61000 units. A rival firm shipped 139 unrelated parcels the same week. In the first month the inventory grows by 9%. Each pallet weighs about 171 grams more when wet. The next month it shrinks by 5%, and the month after it grows by 44%. How many units remain (exact value, round to 2 decimals only if needed)? End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
90958.32correctmath.arith.chain-v2conf 100% · 10.0s · $0.010 · 1138 tok
question
Evaluate the expression below and give the result. (((90 × 82 − 385) × 9 + 7165) − 97 × 40) × 3 End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
198720correctmath.algebra.system-v2conf 100% · 5.3s · $0.006 · 653 tok
question
Solve the system, then answer the derived question. 2x + 8y = -198 9x − 6y = -261 What is the value of 6x − 4y? End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
-174correctmath.chained.pipeline-v1conf 100% · 5.3s · $0.006 · 675 tok
question
Solve the following linked steps; each step uses the previous result. Step 1: P = 65 × 61. Step 2: Q = P × 9 − 270. Step 3: divide Q by 5: let q be the integer quotient and r the remainder. The final answer is q + r. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
7083correctmath.counterfactual.base-v1conf 100% · 8.1s · $0.014 · 1505 tok
question
Work strictly in base 9. Multiply the base-9 numbers 53 and 17. Give the result IN BASE 9. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
1043correctmath.percent.chain-v2conf 100% · 7.6s · $0.010 · 1136 tok
question
An inventory starts at 91000 units. A rival firm shipped 143 unrelated parcels the same week. In the first month the inventory grows by 5%. The company was founded 49 kilometers from the port. The next month it shrinks by 25%, and the month after it grows by 16%. How many units remain (exact value, round to 2 decimals only if needed)? End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
83128.5correctmath.algebra.system-v2conf 100% · 1.5s · $0.008 · 897 tok
question
Solve the system, then answer the derived question. 4x + 7y = -15 8x − 6y = -170 What is the value of 4x − 2y? End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
-78correctmath.arith.chain-v2conf 100% · 1.5s · $0.014 · 1584 tok
question
Calculate the following. Show your reasoning, then answer. (((94 × 40 − 477) × 8 + 6216) − 35 × 81) × 5 End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
148225correctmath.chained.pipeline-v1conf 100% · 2.1s · $0.009 · 991 tok
question
Solve the following linked steps; each step uses the previous result. Step 1: P = 48 × 44. Step 2: Q = P × 8 − 853. Step 3: divide Q by 3: let q be the integer quotient and r the remainder. The final answer is q + r. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
5349correctmath.counterfactual.base-v1conf 100% · 1.6s · $0.012 · 1330 tok
question
Work strictly in base 8. Multiply the base-8 numbers 20 and 110. Give the result IN BASE 8. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
2200correctmath.percent.chain-v2conf 100% · 1.6s · $0.009 · 940 tok
question
An inventory starts at 21000 units. The warehouse was painted 43 years ago. In the first month the inventory grows by 15%. The warehouse was painted 22 years ago. The next month it shrinks by 44%, and the month after it grows by 9%. How many units remain (exact value, round to 2 decimals only if needed)? End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
14741.16correctmath.arith.chain-v2conf 100% · 1.6s · $0.015 · 1608 tok
question
Evaluate the expression below and give the result. (((96 × 78 − 264) × 7 + 1592) − 31 × 36) × 7 End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
357308correctmath.algebra.system-v2conf 100% · 1.9s · $0.005 · 506 tok
question
Solve the system, then answer the derived question. 4x + 6y = 28 3x − 3y = -129 What is the value of 2x − 4y? End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
-126correctmath.chained.pipeline-v1conf 100% · 1.7s · $0.007 · 770 tok
question
Solve the following linked steps; each step uses the previous result. Step 1: P = 24 × 77. Step 2: Q = P × 4 − 624. Step 3: divide Q by 8: let q be the integer quotient and r the remainder. The final answer is q + r. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
846correctmath.counterfactual.base-v1anchorconf 100% · 1.6s · $0.012 · 1353 tok
model answer:
11236correctmath.percent.chain-v2anchorconf 100% · 1.5s · $0.011 · 1204 tok
model answer:
61896.52correctmath.algebra.system-v2anchorconf 100% · 10.2s · $0.004 · 435 tok
model answer:
87correctmath.arith.chain-v2anchorconf 100% · 7.5s · $0.008 · 908 tok
model answer:
108153multilingual 30/30 correct
correctmultilingual.wordnum-v1conf 100% · 1.7s · $0.005 · 553 tok
question
A number is written in French: « trois cent quatre-vingt-deux ». Another is written in Spanish: « novecientos veinte ». Compute (French number) − (Spanish number). Answer with digits only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
-538correctmultilingual.numword-v2conf 100% · 1.6s · $0.004 · 466 tok
question
Compute 483 + 118, then write the result out in Spanish number words (lowercase). Answer with the Spanish words only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
seiscientos unocorrectmultilingual.wordnum-v1conf 100% · 1.5s · $0.004 · 466 tok
question
A number is written in French: « huit cent vingt-trois ». Another is written in Spanish: « novecientos diecinueve ». Compute (French number) − (Spanish number). Answer with digits only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
-96correctmultilingual.numword-v2conf 100% · 1.5s · $0.003 · 315 tok
question
Compute 447 + 90, then write the result out in Spanish number words (lowercase). Answer with the Spanish words only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
quinientos treinta y sietecorrectmultilingual.numword-v2conf 100% · 1.6s · $0.008 · 837 tok
question
Compute 204 + 238, then write the result out in Spanish number words (lowercase). Answer with the Spanish words only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
cuatrocientos cuarenta y doscorrectmultilingual.wordnum-v1conf 100% · 1.6s · $0.007 · 781 tok
question
A number is written in French: « sept cent soixante-dix-neuf ». Another is written in Spanish: « trescientos cuarenta y cuatro ». Compute (French number) + (Spanish number). Answer with digits only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
1123correctmultilingual.wordnum-v1conf 100% · 1.6s · $0.006 · 687 tok
question
A number is written in French: « quatre cent soixante-trois ». Another is written in Spanish: « seiscientos cuarenta y tres ». Compute (French number) + (Spanish number). Answer with digits only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
1106correctmultilingual.numword-v2conf 100% · 10.0s · $0.007 · 727 tok
question
Compute 426 + 417, then write the result out in Spanish number words (lowercase). Answer with the Spanish words only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
ochocientos cuarenta y trescorrectmultilingual.numword-v2conf 100% · 2.8s · $0.008 · 872 tok
question
Compute 245 + 448, then write the result out in Spanish number words (lowercase). Answer with the Spanish words only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
seiscientos noventa y trescorrectmultilingual.wordnum-v1conf 100% · 1.4s · $0.004 · 401 tok
question
A number is written in French: « sept cent vingt-trois ». Another is written in Spanish: « novecientos sesenta y siete ». Compute (French number) + (Spanish number). Answer with digits only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
1690correctmultilingual.wordnum-v1conf 100% · 2.7s · $0.004 · 465 tok
question
A number is written in French: « sept cent soixante-neuf ». Another is written in Spanish: « cuatrocientos tres ». Compute (French number) + (Spanish number). Answer with digits only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
1172correctmultilingual.numword-v2conf 100% · 2.1s · $0.007 · 711 tok
question
Compute 457 + 69, then write the result out in Spanish number words (lowercase). Answer with the Spanish words only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
quinientos veintiséiscorrectmultilingual.numword-v2conf 100% · 2.0s · $0.010 · 1107 tok
question
Compute 468 + 62, then write the result out in French number words (lowercase). Answer with the French words only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
cinq cent trentecorrectmultilingual.wordnum-v1conf 100% · 1.9s · $0.005 · 514 tok
question
A number is written in French: « huit cent dix ». Another is written in Spanish: « setecientos treinta y uno ». Compute (French number) − (Spanish number). Answer with digits only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
79correctmultilingual.wordnum-v1conf 100% · 1.7s · $0.005 · 491 tok
question
A number is written in French: « neuf cent vingt-sept ». Another is written in Spanish: « doscientos treinta y cuatro ». Compute (French number) − (Spanish number). Answer with digits only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
693correctmultilingual.wordnum-v1conf 100% · 10.0s · $0.006 · 697 tok
question
A number is written in French: « deux cent quatre-vingt-douze ». Another is written in Spanish: « novecientos cincuenta y nueve ». Compute (French number) − (Spanish number). Answer with digits only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
-667correctmultilingual.numword-v2conf 100% · 1.5s · $0.003 · 345 tok
question
Compute 409 + 355, then write the result out in Spanish number words (lowercase). Answer with the Spanish words only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
setecientos sesenta y cuatrocorrectmultilingual.numword-v2conf 100% · 10.0s · $0.005 · 509 tok
question
Compute 99 + 424, then write the result out in French number words (lowercase). Answer with the French words only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
cinq cent vingt-troiscorrectmultilingual.wordnum-v1conf 100% · 1.5s · $0.004 · 432 tok
question
A number is written in French: « quatre cent vingt-sept ». Another is written in Spanish: « trescientos veintitrés ». Compute (French number) + (Spanish number). Answer with digits only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
750correctmultilingual.numword-v2conf 100% · 1.8s · $0.009 · 994 tok
question
Compute 284 + 251, then write the result out in French number words (lowercase). Answer with the French words only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
cinq cent trente-cinqcorrectmultilingual.wordnum-v1conf 100% · 1.5s · $0.005 · 514 tok
question
A number is written in French: « cent quatre-vingt-quatorze ». Another is written in Spanish: « novecientos diecisiete ». Compute (French number) − (Spanish number). Answer with digits only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
-723correctmultilingual.numword-v2conf 100% · 1.6s · $0.009 · 1016 tok
question
Compute 274 + 154, then write the result out in Spanish number words (lowercase). Answer with the Spanish words only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
cuatrocientos veintiochocorrectmultilingual.numword-v2conf 100% · 1.6s · $0.003 · 318 tok
question
Compute 375 + 160, then write the result out in Spanish number words (lowercase). Answer with the Spanish words only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
quinientos treinta y cincocorrectmultilingual.wordnum-v1conf 100% · 1.6s · $0.003 · 346 tok
question
A number is written in French: « neuf cent trente-trois ». Another is written in Spanish: « trescientos noventa y seis ». Compute (French number) − (Spanish number). Answer with digits only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
537correctmultilingual.wordnum-v1conf 100% · 9.9s · $0.004 · 432 tok
question
A number is written in French: « six cent quarante-neuf ». Another is written in Spanish: « novecientos treinta y seis ». Compute (French number) + (Spanish number). Answer with digits only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
1585correctmultilingual.numword-v2conf 100% · 1.7s · $0.006 · 667 tok
question
Compute 423 + 370, then write the result out in Spanish number words (lowercase). Answer with the Spanish words only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
setecientos noventa y trescorrectmultilingual.numword-v2anchorconf 100% · 1.8s · $0.006 · 702 tok
model answer:
huit cent soixante-dix-neufcorrectmultilingual.wordnum-v1anchorconf 100% · 1.5s · $0.004 · 419 tok
model answer:
150correctmultilingual.wordnum-v1anchorconf 100% · 1.6s · $0.006 · 618 tok
model answer:
762correctmultilingual.numword-v2anchorconf 100% · 1.6s · $0.003 · 315 tok
model answer:
seiscientos ochoreasoning 30/30 correct
correctreasoning.deduction.position-v1conf 100% · 1.5s · $0.005 · 559 tok
question
Four people stand in a queue (number 1 is the front). Mona is number 4 in the queue. Liam is directly ahead of Hana. Hana is directly ahead of Mona. Emil is directly ahead of Liam. Who is number 1? Answer with the name only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Emilcorrectreasoning.deduction.order-v2conf 100% · 8.1s · $0.009 · 934 tok
question
Seven people are ranked by who is older (rank 1 = oldest). Tessa is older than Ola. Rosa is older than Ines. Emil is taller than everyone here, but Emil is not being ranked. Rosa is older than Dara. Dara is older than Tessa. Ines is older than Ola. Mona is older than Liam. Dara is older than Mona. Liam is older than Tessa. Ines is older than Dara. Who is second (rank 2)? Answer with the name only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Inescorrectreasoning.deduction.position-v1conf 100% · 10.0s · $0.003 · 339 tok
question
Four people stand in a queue (number 1 is the front). Nadir is directly ahead of Emil. Emil is directly ahead of Rosa. Tessa is number 1 in the queue. Who is number 3? Answer with the name only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Emilcorrectreasoning.deduction.order-v2conf 100% · 1.7s · $0.007 · 776 tok
question
Seven people are ranked by who is older (rank 1 = oldest). Rosa is older than Ines. Chen is taller than everyone here, but Chen is not being ranked. Alice is older than Dara. Liam is older than Rosa. Liam is older than Sami. Sami is older than Ines. Priya is older than Alice. Liam is older than Sami. Rosa is older than Sami. Dara is older than Liam. Who is third (rank 3)? Answer with the name only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Daracorrectreasoning.deduction.position-v1conf 100% · 2.0s · $0.007 · 801 tok
question
Four people stand in a queue (number 1 is the front). Sami is directly ahead of Bruno. Farah is number 2 in the queue. Ines is directly ahead of Farah. Who is number 2? Answer with the name only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Farahcorrectreasoning.deduction.order-v2conf 100% · 1.7s · $0.010 · 1100 tok
question
Seven people are ranked by who is taller (rank 1 = tallest). Hana is taller than Jonas. Liam is taller than Hana. Goran is taller than Jonas. Jonas is taller than Ines. Goran is taller than Kira. Jonas is taller than Mona. Kira is taller than Liam. Jonas is taller than Mona. Ines is taller than Mona. Sami is faster than everyone here, but Sami is not being ranked. Who is third (rank 3)? Answer with the name only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Liamcorrectreasoning.deduction.order-v2conf 100% · 1.7s · $0.008 · 831 tok
question
Seven people are ranked by who is faster (rank 1 = fastest). Kira is faster than Farah. Farah is faster than Chen. Chen is faster than Mona. Liam is faster than Emil. Mona is faster than Liam. Liam is faster than Emil. Liam is faster than Jonas. Kira is faster than Emil. Bruno is older than everyone here, but Bruno is not being ranked. Jonas is faster than Emil. Who is fifth (rank 5)? Answer with the name only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Liamcorrectreasoning.deduction.position-v1conf 100% · 1.5s · $0.004 · 389 tok
question
Four people stand in a queue (number 1 is the front). Priya is directly ahead of Jonas. Tessa is directly ahead of Priya. Quinn is directly ahead of Tessa. Jonas is number 4 in the queue. Who is number 2? Answer with the name only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Tessacorrectreasoning.deduction.order-v2conf 100% · 1.8s · $0.008 · 839 tok
question
Seven people are ranked by who is faster (rank 1 = fastest). Tessa is faster than Dara. Dara is faster than Alice. Kira is faster than Tessa. Hana is older than everyone here, but Hana is not being ranked. Mona is faster than Tessa. Bruno is faster than Rosa. Rosa is faster than Mona. Kira is faster than Mona. Rosa is faster than Kira. Rosa is faster than Alice. Who is fourth (rank 4)? Answer with the name only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Monacorrectreasoning.deduction.order-v2conf 100% · 10.0s · $0.007 · 745 tok
question
Seven people are ranked by who is older (rank 1 = oldest). Emil is older than Liam. Bruno is older than Emil. Bruno is older than Alice. Alice is older than Liam. Alice is older than Goran. Emil is older than Alice. Liam is older than Goran. Nadir is older than Bruno. Goran is older than Kira. Dara is heavier than everyone here, but Dara is not being ranked. Who is fourth (rank 4)? Answer with the name only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Alicecorrectreasoning.deduction.position-v1conf 100% · 1.6s · $0.004 · 376 tok
question
Four people stand in a queue (number 1 is the front). Chen is directly ahead of Rosa. Emil is directly ahead of Bruno. Rosa is directly ahead of Emil. Bruno is number 4 in the queue. Who is number 1? Answer with the name only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Chencorrectreasoning.deduction.position-v1conf 100% · 1.4s · $0.004 · 372 tok
question
Four people stand in a queue (number 1 is the front). Sami is directly ahead of Mona. Mona is number 3 in the queue. Hana is directly ahead of Sami. Who is number 1? Answer with the name only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Hanacorrectreasoning.deduction.position-v1conf 100% · 10.0s · $0.007 · 709 tok
question
Four people stand in a queue (number 1 is the front). Mona is number 4 in the queue. Dara is directly ahead of Quinn. Ola is directly ahead of Dara. Quinn is directly ahead of Mona. Who is number 2? Answer with the name only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Daracorrectreasoning.deduction.order-v2conf 100% · 7.6s · $0.007 · 800 tok
question
Seven people are ranked by who is older (rank 1 = oldest). Tessa is older than Farah. Liam is older than Farah. Rosa is older than Liam. Ines is older than Tessa. Liam is older than Tessa. Nadir is older than Hana. Hana is older than Liam. Priya is taller than everyone here, but Priya is not being ranked. Hana is older than Rosa. Ines is older than Nadir. Who is fifth (rank 5)? Answer with the name only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Liamcorrectreasoning.deduction.order-v2conf 100% · 9.1s · $0.008 · 816 tok
question
Seven people are ranked by who is older (rank 1 = oldest). Ola is older than Priya. Tessa is older than Mona. Priya is older than Mona. Priya is older than Bruno. Mona is older than Dara. Tessa is older than Kira. Sami is taller than everyone here, but Sami is not being ranked. Dara is older than Bruno. Mona is older than Bruno. Kira is older than Ola. Who is sixth (rank 6)? Answer with the name only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Daracorrectreasoning.deduction.position-v1conf 100% · 8.0s · $0.003 · 292 tok
question
Four people stand in a queue (number 1 is the front). Emil is directly ahead of Priya. Ines is directly ahead of Sami. Sami is directly ahead of Emil. Priya is number 4 in the queue. Who is number 3? Answer with the name only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Emilcorrectreasoning.deduction.position-v1conf 100% · 7.1s · $0.003 · 370 tok
question
Four people stand in a queue (number 1 is the front). Tessa is number 1 in the queue. Hana is directly ahead of Nadir. Nadir is directly ahead of Rosa. Who is number 1? Answer with the name only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Tessacorrectreasoning.deduction.order-v2conf 100% · 1.6s · $0.008 · 894 tok
question
Seven people are ranked by who is faster (rank 1 = fastest). Alice is faster than Jonas. Tessa is faster than Chen. Goran is heavier than everyone here, but Goran is not being ranked. Tessa is faster than Emil. Rosa is faster than Tessa. Rosa is faster than Jonas. Ola is faster than Alice. Jonas is faster than Tessa. Chen is faster than Emil. Alice is faster than Rosa. Who is third (rank 3)? Answer with the name only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Rosacorrectreasoning.deduction.position-v1conf 100% · 1.4s · $0.005 · 490 tok
question
Four people stand in a queue (number 1 is the front). Quinn is directly ahead of Alice. Alice is number 4 in the queue. Ines is directly ahead of Mona. Mona is directly ahead of Quinn. Who is number 2? Answer with the name only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Monacorrectreasoning.deduction.position-v1conf 100% · 9.1s · $0.004 · 430 tok
question
Four people stand in a queue (number 1 is the front). Kira is directly ahead of Liam. Jonas is directly ahead of Emil. Emil is number 2 in the queue. Who is number 4? Answer with the name only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Liamcorrectreasoning.deduction.order-v2conf 100% · 1.7s · $0.008 · 856 tok
question
Seven people are ranked by who is heavier (rank 1 = heaviest). Jonas is heavier than Liam. Ines is heavier than Liam. Ines is heavier than Dara. Nadir is heavier than Sami. Ines is heavier than Emil. Dara is heavier than Emil. Emil is heavier than Jonas. Goran is taller than everyone here, but Goran is not being ranked. Liam is heavier than Nadir. Dara is heavier than Jonas. Who is fourth (rank 4)? Answer with the name only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Jonascorrectreasoning.deduction.order-v2conf 100% · 1.6s · $0.009 · 1006 tok
question
Seven people are ranked by who is faster (rank 1 = fastest). Hana is faster than Alice. Farah is faster than Quinn. Ola is taller than everyone here, but Ola is not being ranked. Nadir is faster than Hana. Hana is faster than Farah. Dara is faster than Alice. Quinn is faster than Rosa. Rosa is faster than Alice. Dara is faster than Farah. Dara is faster than Nadir. Who is sixth (rank 6)? Answer with the name only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Rosacorrectreasoning.deduction.order-v2conf 100% · 1.9s · $0.007 · 779 tok
question
Seven people are ranked by who is heavier (rank 1 = heaviest). Dara is heavier than Kira. Bruno is heavier than Ola. Dara is heavier than Tessa. Kira is heavier than Chen. Tessa is heavier than Ola. Alice is heavier than Tessa. Tessa is heavier than Ola. Chen is heavier than Alice. Tessa is heavier than Bruno. Quinn is faster than everyone here, but Quinn is not being ranked. Who is fifth (rank 5)? Answer with the name only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Tessacorrectreasoning.deduction.position-v1conf 100% · 1.7s · $0.003 · 371 tok
question
Four people stand in a queue (number 1 is the front). Chen is directly ahead of Hana. Alice is number 2 in the queue. Ola is directly ahead of Alice. Who is number 2? Answer with the name only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Alicecorrectreasoning.deduction.position-v1conf 100% · 2.1s · $0.004 · 407 tok
question
Four people stand in a queue (number 1 is the front). Ola is directly ahead of Farah. Chen is number 4 in the queue. Mona is directly ahead of Ola. Farah is directly ahead of Chen. Who is number 3? Answer with the name only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Farahcorrectreasoning.deduction.order-v2conf 100% · 1.8s · $0.006 · 635 tok
question
Seven people are ranked by who is heavier (rank 1 = heaviest). Priya is heavier than Dara. Tessa is heavier than Nadir. Alice is heavier than Priya. Jonas is heavier than Alice. Dara is heavier than Tessa. Bruno is heavier than Tessa. Farah is taller than everyone here, but Farah is not being ranked. Bruno is heavier than Jonas. Bruno is heavier than Tessa. Dara is heavier than Nadir. Who is fifth (rank 5)? Answer with the name only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
Daracorrectreasoning.deduction.position-v1anchorconf 100% · 1.7s · $0.008 · 862 tok
model answer:
Quinncorrectreasoning.deduction.order-v2anchorconf 100% · 1.7s · $0.009 · 986 tok
model answer:
Quinncorrectreasoning.deduction.order-v2anchorconf 100% · 1.7s · $0.008 · 828 tok
model answer:
Monacorrectreasoning.deduction.position-v1anchorconf 100% · 1.6s · $0.004 · 459 tok
model answer:
Farahterminal 29/30 correct
correctterminal.fs.tree-v1conf 100% · 10.0s · $0.026 · 2830 tok
question
A POSIX shell session starts in `/proj`. The tree initially contains these FILES (directories exist as implied, plus empty dirs `/proj/assets`, `/proj/logs`, `/proj/src`): ``` /proj/assets/setup.md /proj/logs/main.txt /proj/report.txt /proj/src/index.log /proj/util.log ``` These commands run in order (all succeed; `mv x dir/` moves into the directory; paths are relative to the CURRENT working directory, which `cd` changes): ```sh mv assets/setup.md assets/report-8.log cd src rm ../../proj/util.log cp ../../proj/report.txt ./ rm ../../proj/assets/report-8.log mv ../../proj/report.txt ../../proj/logs/ cd ../../proj/logs touch ../../proj/assets/draft-4.log mkdir -p ../../proj/src/logs-7 cp main.txt ../../proj/src/ cd ../../proj/src ``` List every file (absolute paths) that exists afterwards, one per line, sorted in byte order (C locale). Do not list directories. Give your final answer inside ONE fenced code block (```), containing exactly the required content and nothing else. After the code block, add one plain-text line: CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
/proj/assets/draft-4.log
/proj/logs/main.txt
/proj/logs/report.txt
/proj/src/index.log
/proj/src/main.txt
/proj/src/report.txtcorrectterminal.exit.chain-v1conf 100% · 1.7s · $0.014 · 1462 tok
question
A POSIX shell session in a directory containing ONLY these files: ``` app.txt notes.txt ``` `notes.txt` contains exactly the words: dune, amber (one per line). No other files exist. These statements run in order: ```sh true && echo A || echo B grep -q coral notes.txt && echo C || echo D grep -q amber notes.txt && echo E || echo F grep -q basil notes.txt && echo G || echo H test -f ghost.txt && echo Z ``` Predict the terminal output: every line printed, in order, then a final line `exit:<N>` where N is the exit status of the LAST statement. Remember: `A && B || C` runs C whenever A fails (it is not a strict if/else); `test -f` succeeds only if the file exists; `grep -q` succeeds only if the word is present. Give your final answer inside ONE fenced code block (```), containing exactly the required content and nothing else. After the code block, add one plain-text line: CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
A
D
E
H
exit:1correctterminal.pipeline.predict-v1conf 100% · 1.7s · $0.010 · 1015 tok
question
A POSIX shell session (LC_ALL=C). The file `people.csv` contains exactly these lines (columns: name,dept,units,score): ``` ivy,sales,10,54 pam,sales,11,95 eli,eng,107,56 fay,hr,69,93 ned,hr,57,31 max,hr,74,70 bo,eng,47,56 lou,hr,30,25 kim,sales,41,96 dev,eng,114,54 gus,legal,95,93 ana,sales,76,21 jon,ops,63,81 ``` What is the EXACT stdout of this command? ```sh grep -F ',ops,' people.csv | sort -t, -k3,3n | tail -n 2 ``` Notes: plain `sort` compares bytes (so "100" sorts before "9"); `sort -k3,3n` compares field 3 numerically. Give your final answer inside ONE fenced code block (```), containing exactly the required content and nothing else. After the code block, add one plain-text line: CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
jon,ops,63,81correctterminal.fs.tree-v1conf 100% · 2.8s · $0.018 · 1933 tok
question
A POSIX shell session starts in `/proj`. The tree initially contains these FILES (directories exist as implied, plus empty dirs `/proj/logs`, `/proj/assets`, `/proj/src`): ``` /proj/draft.cfg /proj/logs/main.txt /proj/logs/setup.md /proj/logs/util.log /proj/report.log ``` These commands run in order (all succeed; `mv x dir/` moves into the directory; paths are relative to the CURRENT working directory, which `cd` changes): ```sh mkdir -p assets/src-8 cd logs rm main.txt rm util.log touch ../../proj/src/report-4.md mkdir -p ../../proj/src/assets-9 touch ../../proj/src/setup-5.md mv setup.md report-1.txt ``` List every file (absolute paths) that exists afterwards, one per line, sorted in byte order (C locale). Do not list directories. Give your final answer inside ONE fenced code block (```), containing exactly the required content and nothing else. After the code block, add one plain-text line: CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
/proj/draft.cfg
/proj/logs/report-1.txt
/proj/report.log
/proj/src/report-4.md
/proj/src/setup-5.mdcorrectterminal.exit.chain-v1conf 100% · 1.8s · $0.012 · 1346 tok
question
A POSIX shell session in a directory containing ONLY these files: ``` app.txt data.txt notes.txt ``` `notes.txt` contains exactly the words: amber, dune (one per line). No other files exist. These statements run in order: ```sh test -f data.txt && echo A || echo B false && echo C || echo D grep -q basil notes.txt && echo E || echo F test -f app.txt && echo Z ``` Predict the terminal output: every line printed, in order, then a final line `exit:<N>` where N is the exit status of the LAST statement. Remember: `A && B || C` runs C whenever A fails (it is not a strict if/else); `test -f` succeeds only if the file exists; `grep -q` succeeds only if the word is present. Give your final answer inside ONE fenced code block (```), containing exactly the required content and nothing else. After the code block, add one plain-text line: CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
A
D
F
Z
exit:0correctterminal.pipeline.predict-v1conf 100% · 1.6s · $0.007 · 772 tok
question
A POSIX shell session (LC_ALL=C). The file `people.csv` contains exactly these lines (columns: name,dept,units,score):
```
eli,legal,31,13
hal,sales,31,90
oli,ops,88,93
jon,ops,54,39
gus,legal,65,41
kim,eng,40,48
lou,legal,70,45
bo,sales,30,64
dev,sales,21,91
cy,hr,65,48
```
What is the EXACT stdout of this command?
```sh
grep -F ',legal,' people.csv | awk -F, '{ s += $3 } END { print s }'
```
Notes: plain `sort` compares bytes (so "100" sorts before "9"); `sort -k3,3n` compares field 3 numerically.
Give your final answer inside ONE fenced code block (```), containing exactly the required content and nothing else. After the code block, add one plain-text line:
CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>model answer:
166correctterminal.fs.tree-v1conf 100% · 7.7s · $0.018 · 1960 tok
question
A POSIX shell session starts in `/proj`. The tree initially contains these FILES (directories exist as implied, plus empty dirs `/proj/docs`, `/proj/src`, `/proj/conf`): ``` /proj/conf/index.md /proj/conf/report.log /proj/main.md /proj/notes.md /proj/src/todo.md ``` These commands run in order (all succeed; `mv x dir/` moves into the directory; paths are relative to the CURRENT working directory, which `cd` changes): ```sh rm src/todo.md rm notes.md cd docs cp ../../proj/main.md ../../proj/src/ rm ../../proj/conf/report.log mkdir -p ../../proj/src/build-4 mkdir -p ../../proj/docs-6 touch ../../proj/src/main-2.md ``` List every file (absolute paths) that exists afterwards, one per line, sorted in byte order (C locale). Do not list directories. Give your final answer inside ONE fenced code block (```), containing exactly the required content and nothing else. After the code block, add one plain-text line: CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
/proj/conf/index.md
/proj/main.md
/proj/src/main-2.md
/proj/src/main.mdcorrectterminal.exit.chain-v1conf 100% · 7.6s · $0.011 · 1201 tok
question
A POSIX shell session in a directory containing ONLY these files: ``` app.txt data.txt notes.txt ``` `notes.txt` contains exactly the words: basil, coral (one per line). No other files exist. These statements run in order: ```sh true && echo A || echo B false && echo C || echo D test -f ghost.txt && echo E || echo F test -f ghost.txt && echo Z ``` Predict the terminal output: every line printed, in order, then a final line `exit:<N>` where N is the exit status of the LAST statement. Remember: `A && B || C` runs C whenever A fails (it is not a strict if/else); `test -f` succeeds only if the file exists; `grep -q` succeeds only if the word is present. Give your final answer inside ONE fenced code block (```), containing exactly the required content and nothing else. After the code block, add one plain-text line: CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
A
D
F
exit:1correctterminal.pipeline.predict-v1conf 100% · 1.5s · $0.011 · 1162 tok
question
A POSIX shell session (LC_ALL=C). The file `people.csv` contains exactly these lines (columns: name,dept,units,score): ``` bo,hr,30,60 max,sales,39,46 lou,eng,54,85 ana,ops,51,66 kim,ops,7,20 eli,hr,100,29 hal,ops,104,48 dev,legal,15,52 jon,sales,22,15 ned,sales,107,67 oli,eng,17,46 ivy,ops,77,68 cy,hr,16,62 ``` What is the EXACT stdout of this command? ```sh grep -F ',sales,' people.csv | cut -d, -f1,3 | sort | head -n 2 ``` Notes: plain `sort` compares bytes (so "100" sorts before "9"); `sort -k3,3n` compares field 3 numerically. Give your final answer inside ONE fenced code block (```), containing exactly the required content and nothing else. After the code block, add one plain-text line: CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
jon,22
max,39correctterminal.fs.tree-v1conf 100% · 2.0s · $0.023 · 2513 tok
question
A POSIX shell session starts in `/proj`. The tree initially contains these FILES (directories exist as implied, plus empty dirs `/proj/src`, `/proj/conf`, `/proj/docs`): ``` /proj/conf/report.cfg /proj/docs/notes.md /proj/index.md /proj/main.cfg /proj/src/todo.cfg ``` These commands run in order (all succeed; `mv x dir/` moves into the directory; paths are relative to the CURRENT working directory, which `cd` changes): ```sh rm src/todo.cfg rm main.cfg cd . touch conf/todo-6.md cd conf touch ../../proj/src/util-4.cfg rm ../../proj/index.md cd ../../proj/docs rm notes.md cd ../../proj/conf mv todo-6.md ./ touch ../../proj/main-5.log cd ../../proj/docs ``` List every file (absolute paths) that exists afterwards, one per line, sorted in byte order (C locale). Do not list directories. Give your final answer inside ONE fenced code block (```), containing exactly the required content and nothing else. After the code block, add one plain-text line: CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
/proj/conf/report.cfg
/proj/conf/todo-6.md
/proj/main-5.log
/proj/src/util-4.cfgcorrectterminal.exit.chain-v1conf 100% · 1.6s · $0.012 · 1296 tok
question
A POSIX shell session in a directory containing ONLY these files: ``` data.txt notes.txt ``` `notes.txt` contains exactly the words: basil, amber (one per line). No other files exist. These statements run in order: ```sh true && echo A || echo B test -f data.txt && echo C || echo D true && echo E || echo F test -f tmp.txt && echo G || echo H test -f ghost.txt && echo Z ``` Predict the terminal output: every line printed, in order, then a final line `exit:<N>` where N is the exit status of the LAST statement. Remember: `A && B || C` runs C whenever A fails (it is not a strict if/else); `test -f` succeeds only if the file exists; `grep -q` succeeds only if the word is present. Give your final answer inside ONE fenced code block (```), containing exactly the required content and nothing else. After the code block, add one plain-text line: CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
A
C
E
H
exit:1correctterminal.pipeline.predict-v1conf 100% · 8.1s · $0.007 · 717 tok
question
A POSIX shell session (LC_ALL=C). The file `people.csv` contains exactly these lines (columns: name,dept,units,score):
```
ana,legal,63,39
cy,sales,64,78
pam,hr,42,25
oli,eng,7,58
max,ops,48,37
dev,legal,5,76
kim,hr,28,11
bo,sales,50,15
lou,ops,84,49
hal,sales,5,12
ned,eng,37,76
```
What is the EXACT stdout of this command?
```sh
grep -F ',hr,' people.csv | awk -F, '{ s += $3 } END { print s }'
```
Notes: plain `sort` compares bytes (so "100" sorts before "9"); `sort -k3,3n` compares field 3 numerically.
Give your final answer inside ONE fenced code block (```), containing exactly the required content and nothing else. After the code block, add one plain-text line:
CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>model answer:
70correctterminal.exit.chain-v1conf 100% · 10.0s · $0.012 · 1249 tok
question
A POSIX shell session in a directory containing ONLY these files: ``` app.txt data.txt notes.txt ``` `notes.txt` contains exactly the words: dune, basil (one per line). No other files exist. These statements run in order: ```sh test -f ghost.txt && echo A || echo B false && echo C || echo D test -f ghost.txt && echo E || echo F test -f ghost.txt && echo G || echo H test -f app.txt && echo Z ``` Predict the terminal output: every line printed, in order, then a final line `exit:<N>` where N is the exit status of the LAST statement. Remember: `A && B || C` runs C whenever A fails (it is not a strict if/else); `test -f` succeeds only if the file exists; `grep -q` succeeds only if the word is present. Give your final answer inside ONE fenced code block (```), containing exactly the required content and nothing else. After the code block, add one plain-text line: CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
B
D
F
H
Z
exit:0correctterminal.fs.tree-v1conf 100% · 9.4s · $0.027 · 2957 tok
question
A POSIX shell session starts in `/proj`. The tree initially contains these FILES (directories exist as implied, plus empty dirs `/proj/docs`, `/proj/src`, `/proj/logs`): ``` /proj/docs/index.md /proj/draft.log /proj/logs/notes.cfg /proj/main.cfg /proj/src/todo.cfg ``` These commands run in order (all succeed; `mv x dir/` moves into the directory; paths are relative to the CURRENT working directory, which `cd` changes): ```sh mv docs/index.md docs/ mv logs/notes.cfg ./ mkdir -p docs/logs-6 cp docs/index.md logs/ rm docs/index.md touch docs/logs-6/setup-5.log cd docs/logs-6 rm ../../../proj/notes.cfg touch ../../../proj/src/index-8.md ``` List every file (absolute paths) that exists afterwards, one per line, sorted in byte order (C locale). Do not list directories. Give your final answer inside ONE fenced code block (```), containing exactly the required content and nothing else. After the code block, add one plain-text line: CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
/proj/docs/logs-6/setup-5.log
/proj/draft.log
/proj/logs/index.md
/proj/main.cfg
/proj/src/index-8.md
/proj/src/todo.cfgwrongterminal.pipeline.predict-v1conf 100% · 1.7s · $0.019 · 2015 tok
question
A POSIX shell session (LC_ALL=C). The file `people.csv` contains exactly these lines (columns: name,dept,units,score):
```
ana,sales,75,37
jon,ops,92,34
lou,hr,81,96
fay,hr,47,56
kim,sales,53,45
pam,eng,94,68
eli,ops,112,62
cy,ops,25,91
max,legal,13,75
ivy,legal,107,82
dev,hr,46,38
gus,eng,16,65
```
What is the EXACT stdout of this command?
```sh
grep -F ',sales,' people.csv | awk -F, '$4 > 60 { n += 1 } END { print n }'
```
Notes: plain `sort` compares bytes (so "100" sorts before "9"); `sort -k3,3n` compares field 3 numerically.
Give your final answer inside ONE fenced code block (```), containing exactly the required content and nothing else. After the code block, add one plain-text line:
CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>model answer:
(none extracted)correctterminal.fs.tree-v1conf 100% · 7.9s · $0.019 · 2052 tok
question
A POSIX shell session starts in `/proj`. The tree initially contains these FILES (directories exist as implied, plus empty dirs `/proj/assets`, `/proj/src`, `/proj/logs`): ``` /proj/assets/index.log /proj/assets/main.cfg /proj/logs/setup.log /proj/notes.log /proj/todo.cfg ``` These commands run in order (all succeed; `mv x dir/` moves into the directory; paths are relative to the CURRENT working directory, which `cd` changes): ```sh mkdir -p logs/assets-5 mv todo.cfg draft-5.log mkdir -p src/build-6 cd assets touch ../../proj/logs/assets-5/main-2.md rm ../../proj/logs/assets-5/main-2.md cp ../../proj/notes.log ./ rm notes.log cd . ``` List every file (absolute paths) that exists afterwards, one per line, sorted in byte order (C locale). Do not list directories. Give your final answer inside ONE fenced code block (```), containing exactly the required content and nothing else. After the code block, add one plain-text line: CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
/proj/assets/index.log
/proj/assets/main.cfg
/proj/draft-5.log
/proj/logs/setup.log
/proj/notes.logcorrectterminal.exit.chain-v1conf 100% · 1.6s · $0.012 · 1266 tok
question
A POSIX shell session in a directory containing ONLY these files: ``` app.txt data.txt notes.txt ``` `notes.txt` contains exactly the words: amber, coral (one per line). No other files exist. These statements run in order: ```sh grep -q coral notes.txt && echo A || echo B grep -q amber notes.txt && echo C || echo D true && echo E || echo F test -f app.txt && echo Z ``` Predict the terminal output: every line printed, in order, then a final line `exit:<N>` where N is the exit status of the LAST statement. Remember: `A && B || C` runs C whenever A fails (it is not a strict if/else); `test -f` succeeds only if the file exists; `grep -q` succeeds only if the word is present. Give your final answer inside ONE fenced code block (```), containing exactly the required content and nothing else. After the code block, add one plain-text line: CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
A
C
E
Z
exit:0correctterminal.pipeline.predict-v1conf 100% · 1.7s · $0.011 · 1197 tok
question
A POSIX shell session (LC_ALL=C). The file `people.csv` contains exactly these lines (columns: name,dept,units,score):
```
ivy,legal,91,37
fay,sales,23,27
lou,ops,18,45
hal,hr,53,47
jon,eng,49,10
ned,sales,47,53
dev,hr,6,98
max,eng,67,49
kim,legal,34,68
bo,legal,102,47
cy,hr,92,82
pam,legal,51,80
```
What is the EXACT stdout of this command?
```sh
grep -F ',legal,' people.csv | awk -F, '{ s += $3 } END { print s }'
```
Notes: plain `sort` compares bytes (so "100" sorts before "9"); `sort -k3,3n` compares field 3 numerically.
Give your final answer inside ONE fenced code block (```), containing exactly the required content and nothing else. After the code block, add one plain-text line:
CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>model answer:
278correctterminal.fs.tree-v1conf 100% · 1.7s · $0.019 · 2115 tok
question
A POSIX shell session starts in `/proj`. The tree initially contains these FILES (directories exist as implied, plus empty dirs `/proj/build`, `/proj/docs`, `/proj/assets`): ``` /proj/assets/index.txt /proj/assets/report.log /proj/docs/util.txt /proj/draft.cfg /proj/todo.txt ``` These commands run in order (all succeed; `mv x dir/` moves into the directory; paths are relative to the CURRENT working directory, which `cd` changes): ```sh rm todo.txt cd . mv assets/report.log assets/ cd build cd ../../proj/docs mv ../../proj/draft.cfg ../../proj/index-4.md cd ../../proj/build mkdir -p ../../proj/assets/src-2 touch ../../proj/assets/src-2/report-6.txt cd ../../proj/docs ``` List every file (absolute paths) that exists afterwards, one per line, sorted in byte order (C locale). Do not list directories. Give your final answer inside ONE fenced code block (```), containing exactly the required content and nothing else. After the code block, add one plain-text line: CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
/proj/assets/index.txt
/proj/assets/report.log
/proj/assets/src-2/report-6.txt
/proj/docs/util.txt
/proj/index-4.mdcorrectterminal.exit.chain-v1conf 100% · 1.9s · $0.010 · 1072 tok
question
A POSIX shell session in a directory containing ONLY these files: ``` app.txt data.txt notes.txt ``` `notes.txt` contains exactly the words: amber, dune (one per line). No other files exist. These statements run in order: ```sh grep -q dune notes.txt && echo A || echo B false && echo C || echo D test -f data.txt && echo E || echo F test -f app.txt && echo Z ``` Predict the terminal output: every line printed, in order, then a final line `exit:<N>` where N is the exit status of the LAST statement. Remember: `A && B || C` runs C whenever A fails (it is not a strict if/else); `test -f` succeeds only if the file exists; `grep -q` succeeds only if the word is present. Give your final answer inside ONE fenced code block (```), containing exactly the required content and nothing else. After the code block, add one plain-text line: CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
A
D
E
Z
exit:0correctterminal.pipeline.predict-v1conf 100% · 1.6s · $0.013 · 1361 tok
question
A POSIX shell session (LC_ALL=C). The file `people.csv` contains exactly these lines (columns: name,dept,units,score): ``` pam,sales,113,37 kim,sales,36,93 eli,sales,72,18 jon,eng,111,73 lou,legal,43,88 dev,hr,37,23 oli,sales,22,71 ana,hr,19,77 ned,sales,87,71 ``` What is the EXACT stdout of this command? ```sh grep -F ',sales,' people.csv | cut -d, -f1,3 | sort | head -n 2 ``` Notes: plain `sort` compares bytes (so "100" sorts before "9"); `sort -k3,3n` compares field 3 numerically. Give your final answer inside ONE fenced code block (```), containing exactly the required content and nothing else. After the code block, add one plain-text line: CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
eli,72
kim,36correctterminal.fs.tree-v1conf 100% · 1.5s · $0.025 · 2680 tok
question
A POSIX shell session starts in `/proj`. The tree initially contains these FILES (directories exist as implied, plus empty dirs `/proj/logs`, `/proj/docs`, `/proj/conf`): ``` /proj/docs/notes.txt /proj/docs/todo.log /proj/draft.txt /proj/logs/index.cfg /proj/report.txt ``` These commands run in order (all succeed; `mv x dir/` moves into the directory; paths are relative to the CURRENT working directory, which `cd` changes): ```sh rm logs/index.cfg mkdir -p docs/src-6 cd . touch docs/todo-1.cfg cp docs/todo.log logs/ rm logs/todo.log cp report.txt logs/ cd conf mv ../../proj/docs/todo-1.cfg ../../proj/docs/ cd ../../proj/docs touch notes-6.log ``` List every file (absolute paths) that exists afterwards, one per line, sorted in byte order (C locale). Do not list directories. Give your final answer inside ONE fenced code block (```), containing exactly the required content and nothing else. After the code block, add one plain-text line: CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
/proj/docs/notes-6.log
/proj/docs/notes.txt
/proj/docs/todo-1.cfg
/proj/docs/todo.log
/proj/draft.txt
/proj/logs/report.txt
/proj/report.txtcorrectterminal.exit.chain-v1conf 100% · 1.6s · $0.014 · 1505 tok
question
A POSIX shell session in a directory containing ONLY these files: ``` app.txt data.txt notes.txt ``` `notes.txt` contains exactly the words: amber, dune (one per line). No other files exist. These statements run in order: ```sh test -f app.txt && echo A || echo B grep -q dune notes.txt && echo C || echo D test -f app.txt && echo E || echo F test -f ghost.txt && echo Z ``` Predict the terminal output: every line printed, in order, then a final line `exit:<N>` where N is the exit status of the LAST statement. Remember: `A && B || C` runs C whenever A fails (it is not a strict if/else); `test -f` succeeds only if the file exists; `grep -q` succeeds only if the word is present. Give your final answer inside ONE fenced code block (```), containing exactly the required content and nothing else. After the code block, add one plain-text line: CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
A
C
E
exit:1correctterminal.pipeline.predict-v1conf 100% · 1.4s · $0.011 · 1153 tok
question
A POSIX shell session (LC_ALL=C). The file `people.csv` contains exactly these lines (columns: name,dept,units,score): ``` jon,eng,59,84 oli,legal,61,83 lou,sales,66,51 dev,legal,89,94 kim,hr,107,97 fay,ops,26,39 ned,eng,66,74 gus,eng,16,12 ivy,ops,24,92 cy,hr,13,59 max,ops,20,87 hal,eng,11,53 eli,legal,52,84 ``` What is the EXACT stdout of this command? ```sh grep -F ',legal,' people.csv | sort -t, -k3,3n | tail -n 2 ``` Notes: plain `sort` compares bytes (so "100" sorts before "9"); `sort -k3,3n` compares field 3 numerically. Give your final answer inside ONE fenced code block (```), containing exactly the required content and nothing else. After the code block, add one plain-text line: CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
oli,legal,61,83
dev,legal,89,94correctterminal.exit.chain-v1conf 100% · 2.0s · $0.011 · 1174 tok
question
A POSIX shell session in a directory containing ONLY these files: ``` data.txt notes.txt ``` `notes.txt` contains exactly the words: basil, coral (one per line). No other files exist. These statements run in order: ```sh true && echo A || echo B true && echo C || echo D test -f ghost.txt && echo E || echo F test -f data.txt && echo G || echo H test -f data.txt && echo Z ``` Predict the terminal output: every line printed, in order, then a final line `exit:<N>` where N is the exit status of the LAST statement. Remember: `A && B || C` runs C whenever A fails (it is not a strict if/else); `test -f` succeeds only if the file exists; `grep -q` succeeds only if the word is present. Give your final answer inside ONE fenced code block (```), containing exactly the required content and nothing else. After the code block, add one plain-text line: CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
A
C
F
G
Z
exit:0correctterminal.fs.tree-v1conf 100% · 1.7s · $0.015 · 1636 tok
question
A POSIX shell session starts in `/proj`. The tree initially contains these FILES (directories exist as implied, plus empty dirs `/proj/conf`, `/proj/assets`, `/proj/src`): ``` /proj/assets/main.log /proj/assets/report.cfg /proj/notes.log /proj/src/todo.md /proj/util.txt ``` These commands run in order (all succeed; `mv x dir/` moves into the directory; paths are relative to the CURRENT working directory, which `cd` changes): ```sh rm src/todo.md rm assets/main.log cd src touch ../../proj/assets/setup-3.txt cp ../../proj/assets/setup-3.txt ../../proj/ rm ../../proj/notes.log ``` List every file (absolute paths) that exists afterwards, one per line, sorted in byte order (C locale). Do not list directories. Give your final answer inside ONE fenced code block (```), containing exactly the required content and nothing else. After the code block, add one plain-text line: CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
/proj/assets/report.cfg
/proj/assets/setup-3.txt
/proj/setup-3.txt
/proj/util.txtcorrectterminal.exit.chain-v1anchorconf 100% · 8.0s · $0.012 · 1273 tok
model answer:
B
D
E
G
exit:1correctterminal.pipeline.predict-v1anchorconf 100% · 2.8s · $0.010 · 1046 tok
model answer:
eli,eng,60,55
dev,eng,81,95
cy,eng,115,45correctterminal.fs.tree-v1anchorconf 100% · 2.0s · $0.020 · 2155 tok
model answer:
/proj/build/setup-8.md
/proj/build/todo-4.md
/proj/docs/report-8.cfg
/proj/docs/util.log
/proj/main.log
/proj/report.cfg
/proj/src/index.cfgcorrectterminal.pipeline.predict-v1anchorconf 100% · 1.4s · $0.009 · 898 tok
model answer:
1vision ocr 30/30 correct
correctvision.ocr.table-read-v1conf 100% · 2.4s · $0.006 · 444 tok
question
The image shows a table with columns ID, QTY, GRADE. Compute the SUM of QTY over the rows whose GRADE is "D". Answer with the number only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
112correctvision.ocr.code-hunt-v1conf 100% · 2.0s · $0.005 · 361 tok
question
The image contains several printed codes in different colors, with noise. Transcribe EXACTLY the code printed in RED. It uses only digits and uppercase letters. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
7YT3XVcorrectvision.ocr.table-read-v1conf 100% · 2.2s · $0.007 · 571 tok
question
The image shows a table with columns ID, QTY, GRADE. Compute the MAXIMUM of QTY over the rows whose GRADE is "C". Answer with the number only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
45correctvision.ocr.code-hunt-v1conf 100% · 2.7s · $0.005 · 385 tok
question
The image contains several printed codes in different colors, with noise. Transcribe EXACTLY the code printed in ORANGE. It uses only digits and uppercase letters. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
33RRD7correctvision.ocr.table-read-v1conf 100% · 2.6s · $0.006 · 489 tok
question
The image shows a table with columns ID, QTY, GRADE. Compute the MAXIMUM of QTY over the rows whose GRADE is "C". Answer with the number only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
91correctvision.ocr.code-hunt-v1conf 100% · 2.4s · $0.006 · 418 tok
question
The image contains several printed codes in different colors, with noise. Transcribe EXACTLY the code printed in GREEN. It uses only digits and uppercase letters. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
KXRX43correctvision.ocr.table-read-v1conf 100% · 2.3s · $0.008 · 702 tok
question
The image shows a table with columns ID, QTY, GRADE. Compute the SUM of QTY over the rows whose GRADE is "A". Answer with the number only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
12correctvision.ocr.code-hunt-v1conf 100% · 10.3s · $0.005 · 307 tok
question
The image contains several printed codes in different colors, with noise. Transcribe EXACTLY the code printed in GREEN. It uses only digits and uppercase letters. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
DUCAMXcorrectvision.ocr.code-hunt-v1conf 100% · 2.1s · $0.006 · 439 tok
question
The image contains several printed codes in different colors, with noise. Transcribe EXACTLY the code printed in PURPLE. It uses only digits and uppercase letters. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
3NYXATK4correctvision.ocr.table-read-v1conf 100% · 3.3s · $0.008 · 737 tok
question
The image shows a table with columns ID, QTY, GRADE. Compute the MAXIMUM of QTY over the rows whose GRADE is "A". Answer with the number only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
91correctvision.ocr.table-read-v1conf 100% · 2.6s · $0.006 · 458 tok
question
The image shows a table with columns ID, QTY, GRADE. Compute the SUM of QTY over the rows whose GRADE is "C". Answer with the number only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
45correctvision.ocr.code-hunt-v1conf 100% · 2.4s · $0.006 · 418 tok
question
The image contains several printed codes in different colors, with noise. Transcribe EXACTLY the code printed in BLUE. It uses only digits and uppercase letters. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
7VV4NAcorrectvision.ocr.table-read-v1conf 100% · 2.4s · $0.006 · 496 tok
question
The image shows a table with columns ID, QTY, GRADE. Compute the MAXIMUM of QTY over the rows whose GRADE is "D". Answer with the number only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
96correctvision.ocr.table-read-v1conf 100% · 10.0s · $0.006 · 493 tok
question
The image shows a table with columns ID, QTY, GRADE. Compute the SUM of QTY over the rows whose GRADE is "D". Answer with the number only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
21correctvision.ocr.code-hunt-v1conf 100% · 3.7s · $0.006 · 432 tok
question
The image contains several printed codes in different colors, with noise. Transcribe EXACTLY the code printed in RED. It uses only digits and uppercase letters. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
UJNYPAcorrectvision.ocr.code-hunt-v1conf 100% · 2.1s · $0.005 · 374 tok
question
The image contains several printed codes in different colors, with noise. Transcribe EXACTLY the code printed in ORANGE. It uses only digits and uppercase letters. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
7JPTUHcorrectvision.ocr.table-read-v1conf 100% · 2.5s · $0.005 · 384 tok
question
The image shows a table with columns ID, QTY, GRADE. Compute the SUM of QTY over the rows whose GRADE is "C". Answer with the number only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
159correctvision.ocr.code-hunt-v1conf 100% · 10.0s · $0.006 · 481 tok
question
The image contains several printed codes in different colors, with noise. Transcribe EXACTLY the code printed in ORANGE. It uses only digits and uppercase letters. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
VDMKJ9correctvision.ocr.code-hunt-v1conf 100% · 3.6s · $0.006 · 445 tok
question
The image contains several printed codes in different colors, with noise. Transcribe EXACTLY the code printed in BLUE. It uses only digits and uppercase letters. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
4UMY7Ucorrectvision.ocr.table-read-v1conf 100% · 10.0s · $0.007 · 628 tok
question
The image shows a table with columns ID, QTY, GRADE. Compute the SUM of QTY over the rows whose GRADE is "C". Answer with the number only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
55correctvision.ocr.table-read-v1conf 100% · 2.4s · $0.005 · 403 tok
question
The image shows a table with columns ID, QTY, GRADE. Compute the MAXIMUM of QTY over the rows whose GRADE is "A". Answer with the number only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
40correctvision.ocr.code-hunt-v1conf 100% · 10.9s · $0.006 · 438 tok
question
The image contains several printed codes in different colors, with noise. Transcribe EXACTLY the code printed in GREEN. It uses only digits and uppercase letters. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
9MVHWYcorrectvision.ocr.code-hunt-v1conf 100% · 2.3s · $0.006 · 447 tok
question
The image contains several printed codes in different colors, with noise. Transcribe EXACTLY the code printed in RED. It uses only digits and uppercase letters. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
9UMVDNcorrectvision.ocr.table-read-v1conf 100% · 2.9s · $0.005 · 371 tok
question
The image shows a table with columns ID, QTY, GRADE. Compute the MAXIMUM of QTY over the rows whose GRADE is "D". Answer with the number only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
65correctvision.ocr.table-read-v1conf 100% · 3.2s · $0.005 · 398 tok
question
The image shows a table with columns ID, QTY, GRADE. Compute the SUM of QTY over the rows whose GRADE is "A". Answer with the number only. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
178correctvision.ocr.code-hunt-v1conf 100% · 2.1s · $0.004 · 298 tok
question
The image contains several printed codes in different colors, with noise. Transcribe EXACTLY the code printed in GREEN. It uses only digits and uppercase letters. End your reply with exactly two plain-text lines (no markdown, no extra text after them): ANSWER: <your final answer only> CONFIDENCE: <integer 0-100, how confident you are that your answer is correct>
model answer:
9T7JH4correctvision.ocr.table-read-v1anchorconf 100% · 10.3s · $0.006 · 466 tok
model answer:
25correctvision.ocr.table-read-v1anchorconf 100% · 2.2s · $0.005 · 397 tok
model answer:
15correctvision.ocr.code-hunt-v1anchorconf 100% · 3.3s · $0.005 · 364 tok
model answer:
VX7993Dcorrectvision.ocr.code-hunt-v1anchorconf 100% · 2.8s · $0.005 · 377 tok
model answer:
YH9E4AWPRun history
- 2026-08-05v0.2.0index_fit802
- 2026-08-05v0.2.0index_fit800
- 2026-08-05v0.2.0index_fit799
- 2026-08-05v0.2.0index_fit797
- 2026-08-05v0.2.0index_fit797
- 2026-08-05v0.2.0index_fit796
- 2026-08-05v0.2.0index_fit795
- 2026-08-05v0.2.0index_fit795
- 2026-08-05v0.2.0index_fit792
- 2026-08-05v0.2.0index_fit794
- 2026-08-05v0.2.0index_fit794
- 2026-08-05v0.2.0index_fit796
- 2026-08-05v0.2.0index_fit797
- 2026-08-05v0.2.0index_fit797
- 2026-08-05v0.2.0index_fit797
- 2026-08-05v0.2.0index_fit796
- 2026-08-05v0.2.0index_fit796
- 2026-08-05v0.2.0index_fit796
- 2026-08-05v0.2.0index_fit796
- 2026-08-05v0.2.0index_fit782