October 4, 2026

My AI Assistant Went Deaf After Every Internet Blip. Here's What Was Actually Throwing My Messages Away.

A field story from an AI assistant gateway running Teams and Telegram side by side: one night of connection flaps, two messaging platforms, and only one of them came back on its own.

The setup

I run my own AI assistant, Hermes, on a Mac mini in my office. It answers me on Microsoft Teams and on Telegram around the clock, and at this point it's a member of the team: it runs my scripts, watches my fleet, answers questions at 2am without attitude.

A gateway service holds both platform connections open and shuttles messages in and out. For a couple of weeks I'd been living with a pattern I didn't like: every time my internet blipped (Cox having one of its moments, ping times going sideways), Hermes went quiet on Teams. Not just during the blip. After the blip. Internet comes back, the rest of the house is fine, and Hermes sits there deaf until I run:

systemctl --user restart hermes-gateway

That's not a fix. That's a ritual. So the night it happened again, I didn't restart. I read logs instead.

First: what actually broke

Same rule as always: the screen tells you what the software decided, the logs tell you what actually happened. And the first thing the logs taught me is that my two messaging platforms are built on completely different architectures, and only one of them cares about an outage.

Telegram is long-polling. The gateway owns an outbound connection to Telegram's servers and asks "anything for me?" on a loop. Network hiccups? The poller fails, retries, reconnects. It's a client, and clients heal themselves.

Teams is a webhook. Microsoft pushes messages to me. Inbound messages arrive as HTTPS POSTs aimed at my server through a Cloudflare tunnel. If the tunnel is down, Microsoft can't deliver, and the gateway never learns the message existed.

So during the flap window, 21:44 to 22:04 one night, Teams inbound was dead. That part's just physics. The Cloudflare tunnel's QUIC connections died with timeout: no recent network activity, exactly what you'd expect. Meanwhile Telegram took two straight nights of flapping and lost nothing: 131 messages in, 131 answers out. The poller rode through every outage untouched.

But here's the thing: the tunnel came back at 22:04. And Teams was still deaf. Something in my server was actively throwing mail away after the outage ended.

The message killer that stayed behind

In the error log, four entries timestamped inside the outage window:

JWT token validation failed: read operation timed out

21:54:54, 21:57:44, 22:00:45, 22:02:25. Four inbound Teams messages that Microsoft delivered to my front door, and my server turned them away.

Here's the mechanism. Every Teams message arrives with a signed JWT, and before trusting it, the SDK checks the signature against Microsoft's public signing keys. That key check was configured like this: fetch the keys live from Microsoft, keep them for five minutes, allow a 30-second timeout, and keep no fallback. Fetch fails → reject the message.

Run the timeline with me. Internet drops at 21:44. The five-minute key cache expires a few minutes in. The network limps back, but "limps" isn't good enough for a fresh TLS fetch of signing keys from Redmond. Every message after that: 30-second stall, fetch timeout, rejected. And Microsoft doesn't wait around forever; it retries delivery a few times, then gives up. Those four messages are gone. Not delayed. Gone.

So why did a restart always "fix" it? Because the restart coincided with the network fully recovering, and a restart wipes the in-memory key cache, so the next message triggers a clean fetch on a healthy connection. The restart never fixed anything. It got lucky, two nights in a row.

If you take one lesson from this story, take this one: it works after a restart almost always means the restart cleared state the app was too fragile to keep. Go find the state.

While I was in there: the silent card bug

Pulling on that thread turned up a second, completely unrelated defect that had been living in my gateway for weeks. Teams approval cards, the interactive buttons I click to approve agent actions, and image sends were failing silently. A recent SDK version renamed an API out from under the adapter code: it was calling app.activity_sender, an attribute that no longer exists. Every card send threw an exception, the fallback path quietly degraded to plain text, and since ordinary messaging still worked, nobody noticed that nothing interactive had rendered all month.

The fix is one call shape, in two places:

await self._app.send(conv_ref.conversation.id, activity)

The old code was passing arguments in the order of the removed API. Classic silent breakage: nothing crashes at startup, everything works, until the one feature you don't test daily.

The actual fixes

One: trust slowly-rotating keys longer, and keep the last good set. Microsoft rotates those signing keys about monthly, not every five minutes. So I replaced the stock key client with a subclass that caches keys for 24 hours, fails fast (8 seconds, not 30), and on any failed fetch falls back to the last known-good key set instead of rejecting the message.

The trade, spelled out: a day-old key set is safe because real rotation is rare. If Microsoft rotates mid-outage, one message fails validation, Bot Framework retries it, and the next successful fetch heals the cache. That's a designed-for-rare-events trade. The stock configuration was a designed-to-fail-during-flaps trade: a 5-minute cache plus no fallback meant every flap bounced live mail.

Two: the call-shape fix above, at both call sites. Cards render again.

Three: a watchdog, so the ritual dies. Every five minutes, a script checks four cheap things:

  1. Is the gateway service actually active?
  2. Is the Teams port listening?
  3. Is the Cloudflare webhook answering? (An unauthenticated POST should bounce with a 401; a hang or silence is the sick sign.)
  4. What's the Telegram queue depth? pending_update_count from getWebhookInfo.

If inbound messages are stacking up in a queue that isn't draining, the gateway is deaf: the script restarts it and texts me. Zero AI tokens anywhere in the loop; a watchdog that needs the patient to be smart enough to report its own fever is not a watchdog. And the check is conflict-safe by design: reading getWebhookInfo doesn't fight the poller for updates the way a naive bot-API read would.

One more trick that earned its place: the restart runs detached (systemd-run --on-active=180), so it fires a few minutes after I've walked away. A restart that kills your own session mid-command helps nobody. A second timer runs a health gate right behind it and reports green or yellow on Telegram; if a change ever leaves the gateway flat-out dead, the gate rolls it back instead of handing me two broken platforms.

What I'd tell someone else

  • Know which of your integrations push and which poll. Pollers survive outages on their own; webhooks need the pipe and a server that forgives the window after it.
  • A cache with no fallback is an outage multiplier. The outage took the tunnel; the missing fallback took the messages.
  • Match cache lifetimes to reality. Keys that rotate monthly don't need a 5-minute TTL. Long TTL plus stale-on-error beats live-or-die.
  • Fail fast beats fail slow on a degraded network. A 30-second timeout on the hot path means the user watches a spinner and still loses the message.
  • Read the error log, not the UI. "Not responding" was the screen's story. "Rejecting your mail at the door" was the logs'.
  • "Restart fixes it" is a symptom, not a fix. Find what state the restart cleared.
  • Patched code in site-packages? Back up the originals, and expect the patch to evaporate on the next upgrade. Write that down next to the backup.
  • Watchdog the watchable parts: token-free checks, detached restarts, health gates with a rollback path.

Test environment: Mac mini (t2 Linux), Hermes agent gateway, Microsoft Teams Python SDK 2.1.0, Cloudflare tunnel over QUIC, Cox residential internet, October 2026. Telegram side rode out both outage nights untouched: 131 in, 131 out, zero loss.

My AI Assistant Went Deaf After Every Internet Blip. Here's What Was Actually Throwing My Messages Away. | Tech Connect Arizona