Hiring for Potential vs. Experience
A calm reflection on balancing experience and potential in engineering hiring. The article focuses on evidence, learning pace, team needs, and the quiet conditions that help a person grow after they join.
Writing
Deep-dives on software architecture and the way source code is structured — written to be understood by beginners, yet useful to teams shipping at scale. Diagrams, real examples, no hand-waving.
A calm reflection on balancing experience and potential in engineering hiring. The article focuses on evidence, learning pace, team needs, and the quiet conditions that help a person grow after they join.
A calm explainer on treating AI answers as claims that need proportionate verification. The article shows how engineers can keep responsibility by testing AI output against evidence, context, and real system behavior before acting on it.
Preparation often looks like nothing from the outside: a quiet note, a rehearsal, a cleaned-up checklist, or one small risk handled before it becomes visible. This reflection looks at why calm outcomes usually come from work people do not see.
A practical companion to the daily standup conversation, focused on the small decisions a team should make before the day starts: what to finish, what to defer, what to escalate, and where help would change the outcome.
When services hold hands through synchronous calls, latency adds up and one slow dependency takes down checkout. A no-hype guide to event-driven architecture: sync vs async, commands vs events, what brokers really promise (at-least-once, not exactly-once), choreography vs orchestration, and exactly when NOT to reach for events.
Splitting code is the easy half — splitting data is where distributed systems humble you. A practical guide to owning data across services: why a shared database is a distributed monolith, the trade from ACID to eventual consistency, the dual-write bug and the outbox that fixes it, sagas with compensating actions, and when CQRS and event sourcing are worth their lifetime cost.
The moment a call leaves your process it can be slow, fail, or happen twice — and that is the normal case, not the exception. The resilience toolkit, explained plainly: why timeouts come first, how retries become a self-inflicted DDoS without backoff and jitter, why idempotency is the price of retrying, and how circuit breakers, bulkheads, graceful degradation, and observability keep one bad dependency from taking down everything.
The database is almost always the first thing to buckle under growth — and the first thing engineers over-engineer in a panic. A no-hype ladder for scaling the data layer: why you measure and add an index before touching hardware, how read replicas exploit the read/write asymmetry (and the replication-lag trap they bring), where caching helps and why invalidation is the hard part, and when you finally reach for partitioning and sharding — the one decision that is genuinely hard to undo.
In a monolith, debugging was almost cosy — one log file, one process, one place the truth lived. Distributed systems quietly took that away: one request now fans out across a dozen services, and when it breaks there is no single log to read. A no-hype guide to seeing your system at scale: the three pillars (metrics, logs, traces) and the question each one answers, why a single propagated trace ID is the highest-leverage habit you can adopt, how SLOs turn reliability into an error budget you can spend, and how to alert on symptoms so on-call doesn't burn out.