Book an evals working session with us

Your agent is already running.That’s not the same as working

neatlogs finds what’s failing, works out why, and hands you a fix.

trusted by teams at

  • VWO + AB Tasty
  • Runway
  • CFO.ai
  • Gobblecube
  • Vorflux
  • Passionfroot
  • CxConnect
  • Ziva.sh
  • Cipherstash
  • Swirls
24 / Silver flowing veilVersion 21 motion · Version 13 material

import yourtraces fromanywhere

connect your current platform, or instrument directly.

add the contextyour traces are missing

connect customer feedback and business context to agent runs.

catch failuresother tools miss

neatlogs monitors every run, surfaces failures, and groups related cases into a single incident.

app.neatlogs.ai / traces / agent-run-284 production
traces Live
05 sep, 10:16:42 am
search
input

Checkout keeps timing out. Please complete order NL-10482 without charging Maya twice.

what happened14 steps · 24.8s
3 detections
  1. checkout_agent.run. workflow
  2. route_request. http
  3. triage_agent. agent
  4. gpt-4.1. model
  5. customer history ignored
    retrieve_recovery_policytool
  6. decide_resolutionagent
  7. lookup_customer. tool
  8. search_knowledge_base × 9. tool. Deduped
poor judgement

I couldn’t complete the checkout. The checkout service returned a 504 after retrieve_recovery_policy timed out.

RunsLive
AllErrorsSlow
  1. checkout_agent19 errors · 24.8s
  2. support_triage3 errors · 12.6s
  3. returns_agent2 errors · 9.8s
  4. catalog_agent1 error · 7.3s
Agent runNL-10482 · Checkout
247 spans19 errors
WaterfallTreeContext
IssuesCheckout timeout
Live issue · 218 affected runs

Checkout requests time out

HTTP 504
Trace evidence
checkout_agent.runrun_284
refund_agentagent
retrieve_recovery_policytool
new context2 sources added
Evaluation NewRefund policy2 review decisions attached
Attached
Slack New# support-escalationsrefund should have been approved
Attached
incidentssupport_agent · timeout
customer history ignored

The checkout agent started failing immediately after the retry policy changed.

how oftenlast 24 hours
218 failing runs
new context2 signals added
Slack · # support-escalationsrefund should have been approved
Linear · SUP-284recovery policy confirmed
related incidents3 matches
credit offered after cancellation
intent mismatchrelated
repeat failures ignored
recovery judgmentrelated
high-risk account not escalated
escalation judgmentrelated
01

import yourtraces fromanywhere

connect your current platform, or instrument directly.

02

add the contextyour traces are missing

connect customer feedback and business context to agent runs.

03

catch failuresother tools miss

neatlogs monitors every run, surfaces failures, and groups related cases into a single incident.

ship a fix thatstays fixed

neatlogs watches production to make sure it holds.
investigation reportsupport agent · 05 Sep
92%
high confidenceevidence and context agree
root causeconfirmed
retry removedcreate_checkoutv2.8
evidence and reasoning12 linked
Two failing spans, three release comparisons, and four safety checks.
recommended fixready
Restore two bounded retries and return GatewayTimeout after the final retry.
verification3 checks
Timeout path reproduced, duplicate charge blocked, and regression suite passed.

evals so simple evennon-devs can run them

describe what good looks like, and let AI evaluate at scale.

everything your agentsneed to keep working

  • ask @neatlogs anything

    do it in plain english, from Slack or the app. get an answer you can verify.

  • investigate from the terminal

    find the root cause through the CLI or MCP. hand your coding agent the exact fix.

  • catch failures automatically

    detect critical events. route them to the right people with impact, severity, and context.

  • build dashboards by asking

    describe what you want to track in plain language. get a view tailored to your needs.

works everywhere yourteam does

frameworks

  • LangChain
  • CrewAI
  • LlamaIndex
  • Vercel AI SDK
  • Pydantic AI
  • Agno
  • DSPy
  • Maestra

coding agents

  • Claude Code
  • Cursor
  • Codex
  • Copilot

business tools

  • Slack
  • Linear
  • GitHub
  • Jira
  • Notion
  • Discord
  • Webhooks
  • Asana

SDKs

  • Python
  • TypeScript
  • OpenTelemetry
  • REST API

start free.scale when your agents do

bring your own model key at no extra cost. enjoy unlimited users on every paid plan.

paid plans include a free month. if you do not upgrade, your workspace moves to free.

free

for trying neatlogs on a real project.

$0forever
$0forever

what’s included

  • 100k spans a month · hard cap
  • 3 investigations a month
  • traces, sessions & replay
  • human evals, unlimited
  • 5 detectors
  • 2 users · 1 project
  • 14-day retention
start free

starter

popular

for teams shipping agents to production.

$99/month
$84/month

billed annually

everything in free, plus

  • 250k spans a month · then $10 per 100k
  • 50 investigations a month, then $25 per 10
  • 5000 AI evaluations a month, then $15 per 2500
  • cloud agents open a PR from any issue
  • 10 trained classifiers · all detector types
  • unlimited users · 5 projects
  • PII redaction, roles & webhooks
  • 30-day retention
first month free

pro

for teams running several agents at scale.

$499/month
$424/month

billed annually

everything in starter, plus

  • 1M spans a month · then $8 per 100k
  • unlimited investigations & evaluations
  • 100 trained classifiers
  • unlimited detectors, projects & prompts
  • audit log & SSO
  • 90-day retention
  • priority support in a shared Slack channel
first month free

enterprise

for security, compliance, and scale.

customannual pricing
customannual pricing

everything in pro, plus

  • custom span volume · negotiated rate
  • custom retention
  • SAML/SCIM & data residency
  • SOC 2 / ISO, DPA & security review
  • VPC or self-hosted ingest
  • dedicated support with an SLA
talk to us

a span is one ingested step of an agent run (an LLM call, a tool call, or a retriever step). a typical trace is 5–20 spans.

fair question01 / 08

can’t I just hand the traces to Claude?

for one small run, yes.

at production scale, traces are large, repetitive, and mostly scaffolding. more context costs more and can actually make the answer worse.

neatlogs compares the right runs, isolates the critical spans, and sends Claude or Cursor only the failure, likely cause, and evidence.

for one small run, yes.

at production scale, traces are large, repetitive, and mostly scaffolding. more context costs more and can actually make the answer worse.

neatlogs compares the right runs, isolates the critical spans, and sends Claude or Cursor only the failure, likely cause, and evidence.

still unsure?

we’ll look at your stack with you.

get in touch

Your agent is already running.Find out if it’s working

connect your production traces in minutes. no credit card required.