Deep dive 3 min read
Cache-first lookups in an event-driven bot
Why independent event tasks and a cache-first read path were the right shape for an alpha Rust bot, and what that architecture costs you later.
Published
Asuka RS is an alpha Rust side project: an event-driven Discord bot with cache-first lookups, localized responses, and support for 31 locale codes. This note is about two decisions in it that generalise well beyond bots, and about what they cost.
One task per event, not one handler for everything
The naive shape for a gateway consumer is a single loop with a big match. It
works until one branch does something slow or panics, and then the loop that
serves every other event is the thing that stopped.
The alternative is to make the receive loop do almost nothing except hand work off:
while let Some(event) = shard.next_event(EventTypeFlags::all()).await {
let Ok(event) = event else { continue };
let ctx = ctx.clone();
tokio::spawn(async move {
if let Err(error) = handle_event(ctx, event).await {
tracing::warn!(?error, "event handler failed");
}
});
}
Two properties fall out of that. Failure is contained: one handler returning an error is logged and the next event is unaffected. And slow work stops blocking fast work, because a task waiting on a database is not holding the receive loop.
The cost is real and worth stating. Unbounded spawn means unbounded
concurrency: a burst produces as many tasks as there are events, and ordering
between events is no longer guaranteed. For an alpha with human-paced traffic
that trade is fine. At larger scale it becomes a bounded worker pool and a
queue, and you should plan for that rather than be surprised by it.
Cache-first reads, and the invalidation you must not skip
Most bot commands need the same small pieces of state (a user’s preferences, a server’s configuration) for nearly every interaction. Reading those from disk each time is wasted work, so the read path checks an in-memory cache first:
async fn user(&self, id: UserId) -> Result<User> {
if let Some(user) = self.users.get(&id).await {
return Ok(user);
}
let user = self.db.find_user(id).await?.unwrap_or_default();
self.users.insert(id, user.clone()).await;
Ok(user)
}
The interesting part is not the read. It is that every write has to invalidate or update the same key, and that is where cache-first designs actually break:
async fn set_locale(&self, id: UserId, locale: Locale) -> Result<()> {
self.db.update_locale(id, locale).await?;
self.users.invalidate(&id).await; // not optional
Ok(())
}
Miss that one line and you get the worst kind of bug: a user changes a setting, the write succeeds, and the bot keeps behaving as if it did not, intermittently, depending on which node or which TTL window they hit. Write the invalidation in the same function as the write, never in a separate “cache layer” someone can forget to call.
A bounded cache with a TTL turns a missed invalidation from permanent corruption into a bounded staleness window. That is a safety net, not a design.
Read path
Where a user lookup actually goes
Localization as a lookup, not a branch
31 locale codes is the number where string formatting stops being incidental. The rule that keeps it manageable: user-facing text is never constructed in handler code. The handler resolves a locale and asks for a key.
let locale = user.locale.unwrap_or(guild.locale);
let message = self.i18n.lookup(&locale, "command.ping.response");
English is the fallback, always. A missing translation degrades to a language the reader may not prefer, which is inconvenient; a missing translation that panics or renders an empty string is a broken bot. Coverage across those 31 codes can be incomplete (that is the honest state of an alpha), and the fallback is what makes incompleteness survivable.
What I would do differently
Independent event tasks were the right concurrency model for the MVP, and they made the next requirement obvious rather than hidden: once you cannot see how many tasks are in flight, you need real operational tooling. An API and a dashboard belong before a broader release, not after.
The general lesson: pick the concurrency model that contains failure first, and treat the observability it demands as part of the same decision, not as something you will add later when it hurts.