AI Deep Analysis Contacts: I Read 3,000 in Plain Words
I built OT1-Pro because my own sales inbox had 3,000 contacts and I had no idea what any of them wanted. 3,000 rows, tags everywhere, half the chats in Egyptian Arabic, follow-ups slipping, me scrolling like an archaeologist. One night I typed the sentence a customer would later type into my own product: "read the 3000 contacts and analyze them". That sentence became a shipped feature — a run that reads up to ~1,000 contacts, writes findings in plain language, caches for 24 hours at 0 credits, and pings the browser when done. This is how it works, what it costs, and the three expensive bugs that shaped it.
If you run sales over WhatsApp and Messenger, the surrounding pain is familiar. Getting connected is its own war — I wrote the full history in Meta App Verification 2026: A Founder's Guide. Getting the AI to follow up well is another, covered in our sales follow-up automation guide and lead follow-up software breakdown. This post is the next layer: once messages flow, how do you understand thousands of contacts without re-reading them at full price every day?
I typed "read the 3000 contacts and analyze them" into my own product
The feature started as a support ticket. A customer had imported roughly 3,000 contacts from two old spreadsheets and three months of Messenger chats, and he wrote to the AI chat inside OT1-Pro: "read the 3000 contacts and analyze them". No menu item, no reports page — just plain language, the way you would ask a sharp sales manager.
My first version did nothing with that sentence — treated it as normal chat and produced a confident generic paragraph about "diversifying outreach". Useless. The data for a real answer sat in his team database: tags, histories, timestamps, notes. The AI just had no bridge between a casual sentence and a cohort-level job.
So I built the bridge. Today, typing "read the 3000 contacts and analyze them", "analyze all my leads", or "audit every conversation from last month" fires a detector, shows a confirmation card, and one click dispatches a background run. Plain-language themes, objection counts, next actions — not a CSV dump. Everything described here is live in production, not a roadmap sketch.
How the detector actually fires (and why the thresholds exist)
The rule is deliberately narrow — easy to trigger by talking normally, nearly impossible to trigger by accident. It fires only when three things hold at once:
- An analysis verb is present. Words like analyze, audit, review, summarize, breakdown, report, insights, themes, patterns. "Read" alone is not enough — "read the 3000 contacts" does not fire, but "read the 3000 contacts and analyze them" does, because the second clause carries the verb.
- A cohort noun is present. Words like contacts, leads, customers, conversations, chats, pipeline, subscribers. This stops "analyze this reply" from launching a 1,000-contact job when the user meant one message.
- A scale signal is present. Either an explicit number of 20 or more ("analyze my 300 contacts"), or a universal quantifier ("all", "every", "entire", "whole"). Under 20 contacts the normal AI reply path handles it inline; at 20 and above the system offers the deep run.
Why 20? I tested 10, 20, and 50 against three months of AI-chat logs. At 10, the detector fired on questions like "analyze these 12 replies" — jobs that finish in seconds inline and should never touch the queue. At 50, real requests like "analyze my 30 expo leads" slipped through and got weak inline answers. Twenty was the sweet spot. The all/every branch covers the most common phrasing, which has no number: "analyze all my contacts".
One more guardrail: the detector only fires in the team's internal AI chat, never in a customer-facing conversation. A customer typing "analyze all your products" to the sales bot will never launch a cohort job.
What happens after detection: modes, scope, and the 1000-contact cap
When detection fires, the user sees a confirmation card — not an instant charge — naming cohort size, mode, and credit cost. Two modes ship:
- customer_themes. Top objections, buying signals, language mix, dead-lead patterns, three next actions. The mode the "read the 3000 contacts" customer wanted.
- agent_audit. Response-time patterns, follow-up gaps, chats that died after a price question, agents that convert versus stall. If customer_themes asks "what do buyers want?", agent_audit asks "where does our process leak?"
Both modes run as queued background jobs — fan out in batches, call the AI chain per batch, reduce into one report. Each run covers up to ~1,000 contacts; a 3,000-cohort analyzes the 1,000 most recently active first, window stated in the header. The cap keeps queue memory flat, the NaraRouter text chain inside budget, and the report readable. Bigger cohorts run one analysis per segment ("my 800 expo leads", then "my 900 website leads") and get better answers than one giant run.
For teams comparing tools, this is where OT1-Pro diverges from pure pipeline products — a classic OT1-Pro vs WATI comparison covers channels and broadcast pricing, while deep cohort analysis bets the value is in re-reading contacts cheaply every week. If objections are your bottleneck, the sibling AI Sales Agent Objection Handling Playbook pairs well with the customer_themes report.
The credit model: DeepAnalysisService charges on dispatch, except cache hits
Here is the billing rule, exactly as shipped: DeepAnalysisService charges the team on dispatch, except on cache hits, which cost 0 credits. Dispatch-time, not completion-time — the AI calls start within seconds of dispatch, so charging at completion would let failed jobs look free while still burning quota.
Concrete math from production: a ~1,000-contact customer_themes run fans out to ~11 AI calls (10 batch + 1 reduce), ~22 credits — roughly $4.40 of a $79/mo Pro plan credit pool. A single user double-clicking confirm used to dispatch two identical runs: 22 wasted calls, ~$8.80 of credit value gone, two identical reports a minute apart. Worse, I measured one session with 4 identical paid retries during a provider outage — 44 wasted calls, ~$17.60 burned for zero new information. Those numbers forced the lock and the cache.
Failed runs do not silently eat credits. If every provider is down and the job throws AiAllProvidersUnavailable, the run is marked failed and the retry path checks the failure record first — no charge for retrying the identical cohort. Quota and outage failures throw specific exceptions the retry policy understands; anything else returns empty rather than charging for garbage.
The 24-hour 0-credit cache: sha256 of the canonical-sorted filter
The cache is the feature customers feel. Analyzed a cohort in the last 24 hours with nothing changed? Re-running it costs 0 credits and returns in under a second:
- Canonicalize the cohort filter. Team id, mode, tags, date window, search string, and contact cap are serialized with keys sorted and values normalized — tags lowercased and sorted, dates truncated to the minute. "Tags: VIP, expo" and "tags: expo, vip" canonicalize identically. Without this, two logically identical requests hash differently and the cache never hits.
- Hash it. The canonical string is hashed with sha256, scoped to team + mode + hash — Team A's "expo leads" never serves Team B, and a customer_themes result never masquerades as agent_audit.
- 24-hour window. Younger than 24 hours is a hit: 0 credits, instant serve. Older is a miss: full charge. The window is short on purpose — cohorts go stale fast, and I would rather recharge for fresh data than let a team act on dead objections.
The canonical-sorting detail cost me a debugging afternoon. The first implementation hashed the raw request array, so key order mattered: the confirmation card built the filter as {tags, dates, mode} while the re-run link built it as {mode, tags, dates}, and the hashes never matched. Cache hit rate was 0% for a week. Sorting the keys before hashing took the hit rate from 0% to ~60% overnight for daily-active teams.
Dollar math on the cache: a team that checks the same "all leads" analysis every morning — fresh run Monday ($4.40 of credit value), cached re-reads Tuesday through Sunday at 0 credits — spends 22 credits for the week instead of 154. On the $79/mo Pro pool, that is the difference between analysis being a daily habit and a luxury.
The emerald pill: what a cached re-run looks like
A cache hit never pretends to be fresh data. The message carries an emerald-green "cached · 0 credits" pill, the original run timestamp, and the cohort definition ("customer_themes · 1,000 contacts · tags: expo, vip · run 2026-10-08 09:14 UTC"). Below sits a "re-run with fresh data" link that re-dispatches with force_fresh=true — full charge, fresh read, new cache row. Stale data is fine when labeled; stale data pretending to be fresh kills trust.
I tested three versions of this UI. Version one had no pill — just the report. Customers assumed it was fresh, then complained the numbers missed yesterday's import. Version two had a gray "cached" note in small text. Nobody read it. Version three — emerald pill plus timestamp plus explicit re-run link — dropped cache complaints to zero. Green means "this saved you money", the timestamp says how old it is, and the link means "fresh is one click away".
The force_fresh=true parameter bypasses both cache lookup and idempotency lock — the user saying "I changed the data, charge me and re-read". Imports, bulk tag changes, and "we just added 200 expo contacts" are the three legitimate triggers in the logs. Everything else should hit the cache.
The bug that made me build the idempotency lock
Before the lock existed, duplicate paid runs were my most embarrassing credit bug. The sequence was always the same: user clicks confirm, the job takes 30-90 seconds, the user sees no instant result, clicks confirm again — or retypes "analyze all my contacts" and confirms the "new" card. Two jobs dispatch, each charges on dispatch, each fans out to ~11 AI calls. That is 11 wasted calls per duplicate run at minimum, 22 credits double-spent, and two near-identical reports that make the customer feel scammed by my billing.
The fix is an idempotency lock keyed on the same team + mode + sha256 hash as the cache, held from confirm-click through dispatch. The second confirm for an identical in-flight cohort does not dispatch — it returns the in-flight job id with a "this analysis is already running" notice. The lock expires after 10 minutes so a crashed job cannot block its cohort forever. Since shipping it, duplicate dispatches are effectively zero, down from ~12 per week.
The lesson: any button that spends money and takes longer than three seconds needs an idempotency story before it ships, not after. The lock was 40 lines of code. The refunds it replaced were hours of my life every month. New teams can try the whole loop free at the OT1-Pro registration page — double-click the confirm card and watch the lock catch it.
The NaraRouter outage that burned 4 identical paid runs
The worst single incident was not user error — it was a provider outage plus my own retry logic. During a NaraRouter brownout in September, the text chain started returning 500s on batch calls. My job retry policy, borrowed from the normal chat-reply path, retried the whole deep-analysis job 4 times. Each retry re-dispatched the fan-out, each fan-out burned calls against the fallback chain, and all 4 retries failed identically when the NaraRouter 30-min global cooldown engaged. Four identical paid runs, zero reports, one furious customer watching credits drain during an outage he did not cause.
Three changes came out of that night. First, deep-analysis jobs now check the global cooldown state before dispatching the fan-out — if the NaraRouter 30-min global cooldown is active, the job waits with a backoff instead of burning calls it knows will fail. Second, batch-level failures retry only the single failed batch, once, against the next model in the chain — never the whole run. Third, a run that fails with AiAllProvidersUnavailable on every batch is marked failed with no further automatic retries, and the team sees "providers unavailable — retry free when service recovers". Same philosophy as a related guardrail: Meta's error 2018278 outside 24h window refuses to spend on a WhatsApp reply the platform will reject.
Since the fix: zero repeat events in six weeks. The cooldown check alone has absorbed two subsequent brownouts — jobs paused, cooldown cleared, runs completed late but correctly, nobody charged for the wait.
Browser notify via Reverb: DeepAnalysisCompleted and the stale-banner trap
A 60-90 second job needs a better signal than "keep refreshing". When a run finishes, the server broadcasts a DeepAnalysisCompleted event over Reverb on the team's private channel. The inbox layout listens and fires a browser Notification — title, mode, cohort size, click-through to the report. Focused tab? It swaps the "running" banner for the report inline.
The failure mode that bit me: the stale banner when the Reverb event never fires. Corporate networks, ad blockers, and one memorable Safari version all silently drop the WebSocket — the event broadcasts correctly but never arrives, and the user stares at "analysis running…" forever while the finished report sits in the database. Two-part fix: every "running" banner polls job status every 15 seconds (worst case a 15-second delay, not infinity), and every banner carries a timestamp — "running since 09:14:22 UTC" — so anything older than 5 minutes reads as suspicious.
The app requests Notification permission only when the user confirms their first run — the moment the value is obvious. At signup it converted at 11%; at confirm-click, 64%.
When deep analysis is worth it on the $79/mo Pro plan
Honest accounting, using the live credit table (October 2026) and the $79/mo Pro plan pool:
| Pattern | Runs / week | Credits | Credit value | Verdict |
|---|---|---|---|---|
| Fresh 1,000-contact run daily, no cache | 7 | 154 | ~$30.80 | Wasteful |
| Fresh Monday + cached re-reads | 1 paid + 6 cached | 22 | ~$4.40 | Recommended |
| Duplicate double-clicks (pre-lock) | 2 paid per intent | 44 | ~$8.80 burned | Fixed by lock |
| 4 paid retries in outage (pre-fix) | 4 paid, 0 reports | 88 | ~$17.60 burned | Fixed by cooldown check |
| Segmented: 2 fresh runs + caches | 2 paid + cached | 44 | ~$8.80 | Best for 2,000+ |
My recommended routine for 1,000-3,000 contacts: Monday, run customer_themes fresh on your hottest segment. Tuesday, run agent_audit fresh on the same segment — the pair shows what buyers want and where your process leaks. Wednesday through Sunday, re-read from cache at 0 credits, force_fresh only after imports or bulk tag changes. Total: 44 credits a week, ~$8.80 of value, two genuinely useful reports plus free re-reads — leaving the bulk of the Pro pool for actual customer conversations, which is where revenue comes from.
Final note: under 20 contacts, skip deep analysis — ask the AI chat inline, free and faster. Deep analysis earns its keep at 20+ and compounds at 500+.
Stop losing the leads you already earned
OT1-Pro runs your follow-up, analysis, and AI replies in one inbox — WhatsApp, Instagram, Messenger, Telegram, and email, in Arabic or English, scored by lead quality, with every AI credit receipted in a transparent ledger. Free plan, no credit card.
Start free → · Sales follow-up automation · Lead follow-up software · Pricing · vs WATI · Talk to the founder on WhatsApp
جاهز للتجربة OT1-Pro?
اربط واتساب وإنستغرام وفيسبوك وتيليجرام مع ذكاء اصطناعي يبيع نيابةً عنك.
ابدأ مجاناً