• Skip to primary navigation
  • Skip to main content
  • Skip to primary sidebar

Dion Almaer

Software, Development, Products

  • @dalmaer
  • LinkedIn
  • Medium
  • RSS
  • Show Search
Hide Search

Programming

Human concurrency

January 6, 2016 Leave a Comment


“There are only two hard things in Computer Science: cache invalidation, naming things, and off-by-one errors.” — Phil Karlton++?

Concurrency is hard. It is a world of trade offs and you often see some junior engineers thinking they have found the one true solution by adding in a caching layer to fix their performance problem.

I have been there. So many of us have. We need to be able to offer predictable and reliable scalability, and to do that we need to minimize bottlenecks. If we get this wrong though it takes more time to manage replication and invalidation and we have just created a potentially worse, and more complex problem.

As hard as it is to architect a solution that maps to product requirements and SLAs, I often think about how the same problem lies at the heart of organizations.

We have come a long way in how we try to scale humans:

We maybe took it a lil too literally?

But we may have a ways to go, even in 2016:

Thanks to Adam Tait and Stuart Argue for the inspiration and these photos 🙂

How can we predictably and reliably scale?

It is also tempting to grasp for the “simple” solution. If you have grown up around silicon valley for example, you think that the answer is talent density and small teams.

You can’t outgrow those pizza teams, you just need to split them up and you are all set!

However, it may not be that simple. This works just fine if you don’t need to communicate or rely on each other. Normally this isn’t the case. If you have team A relying on team B for some infrastructure (which is defined here as “the stuff you don’t want to worry about to get your work done”) then you could be bottlenecked on their velocity. If team B have multiple client teams then they have to prioritize and work out how to solve for that.

You can easily get frustrated here, especially if there is a lack of transparency around how prioritization happens. Do you have a way to offer up head count or resources so you aren’t just sitting there saying “when are you done” but are diving in to help? Is that even possible given the skill-set needed?

That is often the rub. Humans can’t scale in the same way as CPUs as each one of us is so different, and each one of us has very different software onboard 😉

It is so very easy to get into a log jam. Let’s look at a hypothetical example that you may have seen yourself.

Your company has a team that provides an infrastructure as a service. In general you want to use shared resources, especially if they have domain knowledge that you lack, and if by supporting you all of the ships that they support can rise.

It turns out that they want to support you, but they don’t have the resources to jump on this work. Now begins the prioritization game. Can you collectively get support for this work to be done?

There is nothing worse than being the small fry when it comes to these discussions. Let’s say you are a global company, and you are sharing a central payment service. A small country has a payment solution that is The Gold Standard where they are, so it is crucial to get it integrated into the Global Solution. But then a long comes the Big Gorilla. The country that is so large that one compliance feature goes through the ROI analysis and it knocks the small guys solution below the line. Ouch.

Well, screw it, just go off and do a custom build! We can move fast! Talent density ftw! It feels great while you do it, but you may have created a new cache to invalidate. Fast forward a little and you notice that there are now a slew of other groups that have done the same thing. Team members have moved around and some of those systems are hard to keep alive, let alone gaining features.

How will this play out in the long run? It may still make sense to take this path though, and as someone once told me:

“The life blood of a company is momentum.”

This is why scaling companies can be hard. It will be wasteful and you need to be OK with that. It will be messy and you need to deal with that. You can build a framework to help you make the decisions that you need to make as they come up, and then you need to paddle as fast as hell.


p.s. I am looking forward to seeing what comes out of Scaling Teams. There is a dearth of content on engineering management.

YT? Reactive Working

October 12, 2015 Leave a Comment


Reactive is the new hotness. It can help you scale your backend, and keep your front-end complexity sane. It makes you dinner.

I was reading the fantastic introduction by Andre Staltz on the topic and for some reason my mind combined this stream with another thinking through asynchronous communication at work.


We tend to have a blend of synchronous and asynchronous communications, and sometimes we float between the two. I remember seeing a cool demo back in my webOS days where you would float between asynchronous and synchronous talking between parties. Imagine starting a conversation and going back and forth. Then at some point you and your mate happen to be available at the same time so the communication speeds up and feels synchronous. I am actually surprised that voice communications hasn’t allowed for this, other than silos (e.g. Voxer).

The same thing happens with text; whether it be sms, email, or chat. It gets interesting when you add in the social rules of engagement. If someone sends you a “yt?” message on work slack do you feel like you have to answer within 5 mins?

Actors

Now think of yourself as an actor. You are subscribing to a slew of streams right now, and you need to manage their SLA 🙂

At the end of a productive day you feel like you managed your time well, moving grabbing tasks (events), dealing with them, and then putting them back onto another stream. On these great days I feel like I spent the minimum time possible to move the ball (but no less!). Any thread locks (aka meetings) were productive in that decisions were made and the multiple threads (people) were unblocked.

When the ball doesn’t feel like it is moving we tend to solve via meetings. “If we all just get together and hash it out we will be good to go!” This can sometime work, but much of the time it doesn’t. It tends to go best when people are equally up to date, or able to get up to speed with the latest context quickly and have thought through their side of things.

On the flip side, when you are solving by asynchronously nudging GitHub issues, where people each take time to think through their side of the communication, it can be hard to know how things are going. Grabbing one item and “finishing” it feels productive, but nudging 12 items may actually be more impactful even if it doesn’t seem to be the case.

We have all seen the worst case scenario. That bug that now has 3000 comments and hasn’t been fixed for years. It gets nudged, but not actually towards the goal.

Nudging in the right direction

Maybe that is the key. If the work is moving in somewhat the right direction, then keep pushing and playing the multiple games of chess at the same time. If you notice the work is just being shuffled back and forth like a seasoned bureaucrat? Maybe it is OK to join the threads and use that opportunity to make a large change.

Do you track the velocity of your work? The more I learn about how the conscious brain tricks you, the more I feel like I need to be wary of some gut feelings.

Hard work is Good work

For example, it turns out that science has shown us that you need work to be hard to actually learn efficiently and effectively. The problem is that this hard work may make you feel like you aren’t learning as quickly as you would like. These tricks have very bad side effects for us all.

How often have you gone through the following situation: a certification test occurs once a year, so a week before the next one you cram to be able to pass it. The problem? you have tricked the test, but also yourself and people who depend on you. You now have a false sense of confidence of the material. The cramming helped get some knowledge into memory but probably not durable memory.

What should you do instead? Be a life long learner and test yourself as frequently as possible during the year. When the certification test comes up you will be rightfully confident, because you have proved to yourself that you have a certain level of knowledge. You will also be able to use this knowledge in your job all year round. Oh, and it will take you less time this way by using spaced learning!

If you interleave your learning you will also start to make more connections between your work. Maybe this is why I connected reactive programming with reactive working? YT?

“I have never had a fight with my wife!” The Importance of Resilience

August 25, 2015 Leave a Comment


I overheard a conversation that you have probably heard a variant of yourself. A bloke was so proud of the fact that he hadn’t had a serious confrontation with his spouse to date. Ah, the perfect union.

While some joined him in appreciation, I had to hide the real thoughts going through my head:

“Oh crap, he may be lucky and truly have the perfect situation, or when something does come to a boil, they have never practiced the art of disagreement. They have never worked though a tricky situation”

Resilient Software

This reminded me of a similar conversation that I had awhile back, where I heard an admin act so very proud that one machine had been up for over a year. That scared the hell out of me too. It meant that a restart hadn’t been tested in over a year, and can you imagine the magic and cruft that was built up? That is one of the reasons why folks are excited about immutable servers, or at the very least having systems that get built up from scratch.

I am building a new application, and not only should it be mobile first, it should also be offline first. The majority of experiences should probably be architected in the same way. You notice the opposite these days, and often from applications that were built before the mobile revolution. It is often much harder to bolt on offline capability after the fact.

When you build an offline first client you tend to get some great side benefits:

  • If you are working on local data you can keep a responsive UI (as long as you are smart about keeping off the main thread!)
  • You can progressively enhance when online
  • For example, when DuoLingo matches your typed in answer a simple match can occur locally while a more complex match can be kicked off online (if the client is offline).

You can also get into a situation where you are out of sync. The client and server see differing versions of reality.

Shift+Reload

In retrospect we had it lucky in the Web 1.0 days. We could purely server render and the client was a dumb terminal that recreated itself on every request. As browsers got richer caching we needed to give users the nuclear option: Shift+Reload. Not exactly user friendly, but it sure came in handy (and still does!).

These days we need to make sure that our rich clients aren’t getting corrupt. For our clients to work well offline they have local state and data to work on. We have all been frustrated when there is a bug that you can’t easily restore from.

Borken Downloads

One example for me is installing applications on my devices. As I type I have gotten into the situation where my download is hung, yet I have no way to kick start it or even delete it. There is something so very infuriating when this happens, when a version of a shift+reload doesn’t fix the situation. Just yesterday a coworker and I created the same projects on Asana because we didn’t see that the other had already done so. It took forever for me to see his version, and be able to clean it up.

It is tough to get this right. We are trying to do the right thing for the user by caching and keeping their application responsive, yet we should take some hints and have systems to help out. If a user is killing and restarting their application, that is the modern shift+reload is it not?

Micro Services

In theory the birth of micro services and trying to hide the complexity of functionality behind nice clean decoupled APIs helps us with resiliency. In practice I have seen this turn out to be a real mess. It isn’t the fault of the practice, but rather the implementation details. Here is what I have seen go wrong:

The scope of the services isn’t defined correctly

  • In one example the scope seemed to be a function of team size vs. the natural composition of the functionality

Inter-dependency killers

  • There were separate small services, but they all depended on each other. The result was a lot of communication around “a new version of service X was deployed and it broke service Y in QA”

No view of the system as awhole

  • When something goes wrong, how are you made aware? How do you then find where the problem is? Due to not having enough of a view on the whole system it can be hard to get this information. You have to explicitly spend time on the seams

Poor exception handling

  • I hate it when all errors and exceptions are treated as equal. This results in socket closed exceptions being thrown into the mix where they weren’t errors at all…. the client just disconnected and it was fine! As soon as you get this wrong you get a sea of information that you can’t trust, and the killer errors can go un-noticed. I have seen shocking bugs live in production for far too long due to this :/

Finger pointing

  • The worst situations occur when you have constant finger pointing. Something is wrong in the system but each team is arguing about what is actually broken. Services teams point at each other, and point at the network guys, who point at the infrastructure guys who …. point back to the services folk!

Spending time up front to get ahead of this is vital. Certain platforms shine here too. Erlang is known for holding resiliency as its core tenet. Various reactive platforms do well, but although these can make life better for you, you need to care.

I have often had to hold my nose and do the impure. I have setup proxy layers that do automatic retries when the core backend should have been fixed. This is risky, because you can end up increasing the traffic and causing even more issues, but if done right it can save your bacon.

Have you gone through and spec’d the SLA needed for various services? So often we see a least common denominator when it is better to split things out. As an example, if you look at an API that gives you information on a product (description, price, availability, reviews, images, etc) you may want an up to date price but those reviews? Not so much. You can probably deal just fine without that one review that just came in. In this case you probably want to say the equivalent of:

“try to get the latest reviews, but if they don’t come back in Xms then use the last grabbed…. and when that call comes back update the cache for next time maybe, cool?”

It isn’t a surprise that hapi, which Eran Hammer and his team started with me at Walmart, does this pretty well thanks to a box for a cat, as well for handling microservices in general.


For a great modern experience that is fast and works well for your users, chances are that you should:

  • Build an offline first client, but give it enough intelligence to be able to handle corruption and get back to a clean state of health, even with nuclear options
  • Build a services tier that assumes failure at each tier, and that can deal with that failure gracefully
  • Progressively enhance the experience on both the client and the server to make sure the core service always works, but that it can also turn on features and tweaks when available.

And as soon as you have something running, start taking your code to counselling so the system can get good at dealing with disagreements and disruption 😉

« Previous Page

Primary Sidebar

Twitter

My Tweets

Recent Posts

  • Stitching with the new Jules API
  • Pools of Extraction: How I Hack on Software Projects with LLMs
  • Stitch Design Variants: A Picture Really Is Worth a Thousand Words?
  • Stitch Prompt: A CLI for Design Variety
  • Stitch: A Tasteful Idea

Follow

  • LinkedIn
  • Medium
  • RSS
  • Twitter

Tags

3d Touch 2016 Active Recall Adaptive Design Agile AI Native Dev AI Software Design AI Software Development Amazon Echo Android Android Development Apple Application Apps Artificial Intelligence Autocorrect blog Bots Brain Calendar Career Advice Cloud Computing Coding Cognitive Bias Commerce Communication Companies Conference Consciousness Cooking Cricket Cross Platform Deadline Delivery Design Design Systems Desktop Developer Advocacy Developer Experience Developer Platform Developer Productivity Developer Relations Developers Developer Tools Development Distributed Teams Documentation DX Ecosystem Education Energy Engineering Engineering Mangement Entrepreneurship Exercise Eyes Family Fitness Football Founders Future GenAI Gender Equality Google Google Developer Google IO Google Labs Habits Health Hill Climbing HR Integrations JavaScript Jobs Jquery Jules Kids Stories Kotlin Language LASIK Leadership Learning LLMs Lottery Machine Learning Management Messaging Metrics Micro Learning Microservices Microsoft Mobile Mobile App Development Mobile Apps Mobile Web Moving On NPM Open Source Organization Organization Design Pair Programming Paren Parenting Path Performance Platform Platform Thinking Politics Product Design Product Development Productivity Product Management Product Metrics Programming Progress Progressive Enhancement Progressive Web App Project Management Psychology Push Notifications pwa QA Rails React Reactive Remix Remote Working Resilience Ruby on Rails Screentime Self Improvement Service Worker Sharing Economy Shipping Shopify Short Story Silicon Valley Slack Soccer Software Software Development Spaced Repetition Speaking Startup Steve Jobs Stitch Study Teaching Team Building Tech Tech Ecosystems Technical Writing Technology Tools Transportation TV Series Twitter Typescript Uber UI Unknown User Experience User Testing UX vitals Voice Walmart Web Web Components Web Development Web Extensions Web Frameworks Web Performance Web Platform WWDC Yarn

Subscribe via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Archives

  • October 2025
  • September 2025
  • August 2025
  • January 2025
  • December 2024
  • November 2024
  • September 2024
  • May 2024
  • April 2024
  • December 2023
  • October 2023
  • August 2023
  • June 2023
  • May 2023
  • March 2023
  • February 2023
  • January 2023
  • September 2022
  • June 2022
  • May 2022
  • April 2022
  • March 2022
  • February 2022
  • November 2021
  • August 2021
  • July 2021
  • February 2021
  • January 2021
  • May 2020
  • April 2020
  • October 2019
  • August 2019
  • July 2019
  • June 2019
  • April 2019
  • March 2019
  • January 2019
  • October 2018
  • August 2018
  • July 2018
  • May 2018
  • February 2018
  • December 2017
  • November 2017
  • September 2017
  • August 2017
  • July 2017
  • May 2017
  • April 2017
  • March 2017
  • February 2017
  • January 2017
  • December 2016
  • November 2016
  • October 2016
  • September 2016
  • August 2016
  • July 2016
  • June 2016
  • May 2016
  • April 2016
  • March 2016
  • February 2016
  • January 2016
  • December 2015
  • November 2015
  • October 2015
  • September 2015
  • August 2015
  • July 2015
  • June 2015
  • May 2015
  • April 2015
  • March 2015
  • February 2015
  • January 2015
  • December 2014
  • November 2014
  • October 2014
  • September 2014
  • August 2014
  • July 2014
  • June 2014
  • May 2014
  • April 2014
  • March 2014
  • February 2014
  • December 2013
  • November 2013
  • October 2013
  • September 2013
  • August 2013
  • July 2013
  • June 2013
  • May 2013
  • April 2013
  • March 2013
  • February 2013
  • December 2012
  • November 2012
  • October 2012
  • September 2012
  • August 2012

Search

Subscribe

RSS feed RSS - Posts

The right thing to do, is the right thing to do.

The right thing to do, is the right thing to do.

Dion Almaer

Copyright © 2026 · Log in

Loading Comments...