Friday, March 17, 2023

Technical design: whether, who, how, and what

This is a lightly edited update to this post originally published on 20 Aug 2018

Do I need a technical design?


In agile software development, there is architecture (decisions that are hard to change) and incremental design. Architecture, in this sense, is a pretty small number of things—programming language and probably application frameworks and data storage. Incremental design is the norm: we add classes, endpoints, and database tables as we identify a need for them, or remove them as they are unneeded or replaced.

But what about decisions in between these two extremes? For example, it used to be that users all signed up for the website as individuals, and now there is a need for some kind of organization which can manage the users under it. Or we used to have a bunch of separate products with their own logins, apps, and management and now there is a need to do some or all of those things in ways which apply to all products. Or our application used to assume that all users needed to be connected to the internet at all times, and now we want to build in offline operation.

I won’t completely rule out handling larger changes via the usual communication of incremental development—pair programming, discussion of individual stories, pull request review, and the like. But it can be hard to maintain a clear idea of the larger design that way, and I have usually been happier with a discussion which happens at a higher level and whose goal is to get a direction into which we can fit in the smaller decisions that we will make as we go.

I’ll write more later about who should drive this process, how to develop such a design, and what is worth writing down and communicating. But let's first ask when we should be doing this design.

It is tempting to say that the high level design of a system must happen before we can start breaking down the work or implementing pieces of it. Which sounds good, and is nice when it works out, but I have yet to see a design of this sort which does not get changed during implementation. There’s a lot of reality check (interactions with existing functionality, feedback which we only get when we have an early version to show, complications which we didn’t notice at first). Therefore I wouldn’t try to finalize the design before we start acting on it. And I wouldn’t go to the other extreme—of trying to make major changes in a fully incremental way and doing all the communication after the fact. My preference is to start with rough ideas and conversations about the design, and as those get refined and conversations continue, there is a point where the general contours start falling into place. That’s about when I start implementation. I want at least some of the coding to be happening (even if we know we might be revising it later), because otherwise I don’t really trust the design. In parallel, I’m stepping up the communication (documents, meetings, etc). As things fall into place (which may include allocating people’s time, agreeing on technical or business decisions, and getting a clearer picture of implementation choices), you’ll fall into the rhythm of building the thing, because the general contours of what you are building have been established by this point.

Who drives a technical design?


So we have a problem which is large enough that we don’t think we want to approach it in a purely tactical way, and we’ll even assume we have defined at least the general outlines of what we want this design to accomplish. Who should turn this into a design detailed enough to implement?

Before I discuss who, let me say this is an intrinsically messy process. There are a bunch of things we want out of our design. Things to do now or save for another day. People (in various roles) with opinions (either because, well, people have opinions, or more nobly, because they have a specific organizational goal they are trying to achieve). See for example Gregor Hohpe’s The Architect Elevator. Issues like reliability, security, accessibility, and branding. A large design space (a distinguishing character of software being its malleability—or at least potential for malleability). Pros and cons for pretty much every aspect.

If that seems daunting, don’t despair. Just don’t be surprised if a decision which was discussed at length, carefully considered, agreed by all, and signed off subsequently starts to seem less settled. Or someone who you had thought was aware of what was going on suddenly “discovers” your design and has suggestions. Or your scope seems to keep expanding or contracting.

The most important person in this process is the one who is refining the design and who will be involved in implementing it. We can call them the “responsible” person (although don’t think of the roles too rigidly—I did say this process tends to be on the messy side, didn’t I?). To do all these things, and have time for this design, the responsible person needs to be able to focus on this (usually, this means they aren’t a manager).

But that person can’t produce a good design by sitting in a room and thinking hard (if for no other reason, because getting buy-in is a key part of what will make this design get implemented and achieve its goals). Therefore their main activity is going to be communication. I’ll talk later about how to communicate and what to communicate, but in the context of “who”, identify who should be “consulted”. That is, who needs to be aware of the design and would have good ideas about how to do it. Broadcasting what you are doing and inviting input works well, but I’d also directly seek out the people who will be most knowledgeable or important.

One rule of thumb for involving a lot of people is “accept input widely, accept direction narrowly”. You want to hear from as many perspectives as you can. Whether or not you take the advice, thank people and appreciate that they took the time to engage with you. These will be the people who help communicate the changes you are making.

Saying “accept direction narrowly” raises the question of who ultimately will be deciding. This role is generally called the “approver” and will often be the manager of the responsible person (the details will depend on your organization, though). Sign-offs are a good way of formalizing decisions already made and making sure that there is sufficient buy-in throughout the organization. They aren’t good at exploring different possible solutions or weighing pros and cons, so think of formal sign-off type processes (if you have them) as a way of ratifying what is already understood, not as a way of hashing out agreements.

Lastly we have people who aren’t necessarily providing input but who should be “informed” about the design. The basic goal here is to cast as wide a net as feasible (in accordance with “err on the side of overcommunicating” which tends to be good advice especially in larger organizations). Think of ways to reach a variety of audiences: different levels of detail, different ways of presenting the work (for example, it can work to have one document which is technical and one which is more about the business goals and rationales—as long as they are reasonably in sync on topics such as what is in or out of scope), or different places you can announce what you are doing and offer to answer questions or sync up with interested parties.

Describing the responsible, approver, consulted, and informed roles makes it clear that communication is central to the process of making technical decisions and being ready to put them into practice. The next two parts of this series will be about how to communicate, and what topics to include in that communication.

How do I develop and promote my technical design?

In the first two parts of this series we figured out we needed some kind of technical design, and we figured out who should be making that happen. How does the responsible party get this thing going? Do you call a meeting? Write something up?

Typing “useless meeting” into an internet search engine and reading the results should be enough to give us pause about calling a meeting to hash out our technical design. Yet in so many organizations the meeting is the mechanism by which attention is allocated, or is otherwise necessary. So first, what are the pitfalls? The usual risk of a meeting turning into (too much of) an open ended discussion is exacerbated by the large design space and many stakeholders. Another sign that meeting discussion is a bad idea is if the wrong people are there: don’t hesitate to say “can the three of us (less than the whole meeting) have a break-out on this topic after the meeting?” or “would you be willing to talk to X (who is not present) and bring the information back?” Set your goals, such as (1) make a brief announcement about what is underway and how people can get more details or engage further, (2) present your design to date and solicit clarifying questions, or (3) give people an opportunity to raise concerns to be addressed in the future. Or if you do want a longer discussion, set the topic, keep an eye on the clock, and don’t be afraid to steer the group back to the agenda. Also, aim for a level of detail appropriate for the people in the meeting. Software developers may be most interested in database schemas and code organization, infrastructure engineers may be most interested in reliability, security or how your design is spread across various machines, product may be most interested in what functionality your design will or will not unlock, and so on.

I’ve often had good luck circulating the design in document form. People have something to react to and can leave comments on the document itself or in other ways. So is this a Big Design Up Front? Not exactly. I’m aiming for something closer to a High Level Design Written As We Need It. It is at a higher level than code. It is at a higher level than detailed descriptions of functionality (click on button X and see the following fields with the following error conditions). It might contain things like database schemas or protocol specifications, although sometimes even that can be a bit fine grained.

What is a design document for? First of all, as a communication tool. Secondly, to clarify the thinking of the person writing it. What about things like traceability between requirements and implementation, justifying the need for making a change, or documenting what has been changed? I would tend to think of those kinds of documents (how many you need will vary depending on your situation) as separate. The design doc is written and revised as you are thinking something through and figuring it out. More concrete documents (including breakout into tasks, specifying behaviors in detail, or explaining code details), have a greater need for detail and precision and are the output of the design process, although of course the design document can link to them as they are created. Seeing the design document as a communication tool helps focus the process of writing it. Imagine that it is a conference talk and you are trying to figure out who is the audience and what they would want to know about your design.

Expect to iterate on the design. Gather some ideas. Think about them and boil them down to a proposed design. Talk to people one on one. Circulate it in writing. Figure out how else to get it out there. That will generate ideas and reactions. Figure out what to revise based on that. Expect to repeat this process until there is a sufficient degree of convergence on a course of action. Don’t fall into either the extreme of spending all your time talking to people (and not getting around to taking in what they said, researching things as needed, and making some decisions), or the other extreme, of thinking through something and coming up with something which makes sense to you, but which may lack buy-in from other people or may miss important requirements.

So we are developing our design and communicating in diverse ways (presentations, written documents, informal discussions, and yes maybe even meetings). But what topics should we cover? The last section goes into some specifics.

What goes into a technical design?

So far we decided we need a technical design, figured out who would be doing it, and how we’ll be sending it out and getting input. But what is the content of that communication (for example, what sections would we put into a written design document)?

What to include will vary depending on your organization and the needs of a particular design. For an early stage startup, anything relating to scaling and operations may take a back seat to “am I building something people want and how can I most quickly validate my hypothesis?”. For a company in a highly regulated space, there may be a lot of requirements specific to your field.

The same applies to an individual design. Does my design concern a server with a high or low need to be available? Does my design concern data which is sensitive? Does this design change anything related to this topic? (If not there’s probably little to say on the subject). For that reason, I’d suggest treating templates (including this article) as guidelines, and omitting sections which don’t seem relevant. One of the fastest ways to lose an audience is to include a bunch of material that you aren’t very interested in (and probably didn’t do a very good job with). And of course to prioritize everything is to prioritize nothing, a good motto in a variety of contexts.

So, what might we include?

Goals and non-goals
These are perhaps the most important sections. If you can figure out what your design achieves and what you are leaving for another day or deciding is not worth doing, you are well down the path of figuring out how to do it.
Description of the proposed solution
What changes will we make to code, data, networks, and hardware? How does this design achieve the goals? Give enough detail that people can see some of the implications of various choices, but try to avoid the kinds of details which can easily be fleshed out during implementation.
Security
What data is stored and sent where? How is access controlled? If cryptography is involved, how are keys managed and have we chosen appropriate algorithms? Are some parts of the system isolated from others and if so how?
Reliability
Is there redundancy? What are the consequences of network outages? If data is stored in a primary-replica setup, how do we choose a new primary? If data is written multiple places how do we reconcile them? Are there rate limits or other ways of keeping a problem one place from cascading elsewhere?
Capacity
What is the expected load on the various systems involved? Does load ramp up gradually or do we expect a sudden spike in traffic? What needs to be handled manually and is there sufficient staffing to do it?
Monitoring
Do we need to report new metrics? How will we know about errors?
Data analytics
How will we measure usage of the new functionality? What kind of analysis might we want to do?
History
Has the company considered this problem before? What previous decisions got us here? If there are documents describing previous designs, I tend to just link to them rather going into a lot of detail about what has gone before.
Storage
What database(s) are involved (new or existing)? What changes in database schemas are required?
Interfaces between systems
Defining these can help clarify the design and is particularly helpful if one of the functions of your design is to coordinate between different teams or companies who are responsible for different pieces.
Alternatives
How else did we consider solving the problem? Why did we choose the solution we are proposing?
Open questions
This section is particularly helpful if you know certain topics are controversial or warrant further discussion. As questions are resolved, move items from here into the main design section or the alternatives section.
Rollout
In what order are we building this? Are we shipping it continuously? In a series of phases? Is it rolled out selectively to certain users?

These questions can be taken as a template for a design document, but they also can be used to figure out who to go talk to, what to put into a presentation, or what anticipated questions to prepare for.

I've talked a lot about things to do: Did you talk to X? Did you consider Y? What if we did Z? And those are all very helpful up to a point. But only do those things which seem necessary for your particular organizational culture and problem you are trying to solve. The purpose of all these suggestions is to help you build things and solve problems, so as you go, don’t be afraid to keep asking yourself and others: Are people on the same page now? Is this enough specificity to build this? Is my technical design sufficient for what I need?

Monday, June 28, 2021

My company got acquired! Now what?

So you've been working on this product, got something interesting, probably even some customers, and a larger company got interested in the product enough to acquire the company (some of this advice also applies if they were interested in the people or something else other than the existing product, but I'm mostly writing here about the case where - at least supposedly - the intention is to keep and grow the product you were working on before the acquisition).

Some of the following is broken down by function but even more than usual this is a case where it pays to have some cross functional awareness of what is happening even if you are responsible for one of these areas more than others.

First of all, you are now part of a much bigger company. Therefore, a lot of the challenges are communication ones. There's a whole set of issues around getting to know people in disparate parts of the organization, setting expectations, self promotion, and probably a bunch I'm forgetting to call out specifically. But many of them change and get more important at a bigger company.

An acquisition tends to bring up a lot of feelings - for example excitement, accomplishment, sadness, and disorientation. Particularly for people managers, but also everyone, a lot of the job is talking people - including yourself - down from various ledges and getting information about benefits, offices (and/or remote practices), org charts, company strategy, and more. Basically to make sure people know what is going on, have input as feasible, and that things like 1:1s are doing the job of getting into issues which might be less amenable to blanket emails and the like.

For product, the key challenge is "are we working on what the higher ups acquired this product for?" But before even getting to that, regularly ask a more basic question: "Do people widely understand our product (existing and future functionality)?" The big company adage is "err on the side of overcommunicating" and my experience is that you need a lot of different ways to even get to a shared baseline of what we have today (for example, via demo days, screenshots, getting people internally to try the product, and bringing in customer input). Many of the same mechanisms also apply to a shared understanding of what we want to build next and why.

For technical issues, how does the other company handle deployment? Security? Programming languages? Testing? How much do we expect to standardize and how much do we expect to remain divergent? If we want to converge, what do we tackle first and how?

Culturally, the number one thing I'd focus on is how to have contact with people who had been from the other company. Maybe there are interest groups around hobbies, diversity, or charitable activities. Or more work-related things like "people using a common programming language", "people interested in security", or other concerns which may cut across the org chart. It can be easy to neglect things which aren't tied to a concrete deliverable, but the goal here is to build relationships. It is so much easier to navigate an unfamiliar organization and solve a tough problem if you know people and understand assumptions or typical ways of approaching things.

Will you stay with a company for long after acquisition? Does a product have a good change of thriving after acquisition? I'm not diving deeply into that, and there is no shame if the company you end up leaving in a few (months, years, whatever) just doesn't feel like the one you worked for pre-acquisition. But here I try to present the optimistic case for how you can jump into tackling a work situation which just changed (perhaps very dramatically) when your company was acquired.

Sunday, May 09, 2021

The Mortifying Ordeal of Soloing All Day

I read The Mortifying Ordeal of Pairing All Day by Nat Bennett, and.... well first of all it is a good read and worth trying to take in. Perhaps the best way to respond with such a heartfelt and personal story is with my own story. Maybe some day I'll find another way to tell it, but it rang surprisingly true to just take that article and instead lightly edit it to be about my own struggles with soloing (including in many companies where pairing had once been a norm but then fell away for various reasons). The result takes a few liberties here and there but is on the whole autobiographical.

The Mortifying Ordeal of Soloing All Day

I had to confront a lot of my fears about myself, sometimes every day. I had to learn to show someone else all the things I didn’t know, my limitations as a human and a software engineer.

From 2014 to 2020 I was part of an experiment: I soloed all day, most days, for years. Hundreds of other engineers joined me in this experiment. I was working as a software engineer for some of Tech's most exciting startups, and everyone soloed, often for eight hours a day.

This was one of the best things I’ve ever done for myself, socially and emotionally, and it produced some great software. It also burned me out. Not “I don’t want to think about work” burnout. Probably not even “I don’t want to work ever again” burnout. Whether "burnout" is even the exact best word is unclear, but it is close enough to describe a situation where I was often worrying about work (not sleeping well for example).

I spent much of 2020 in discussions with management about how I was doing which led to leaving my job in May 2021 with no plan much more concrete than "give myself and the world a few months to breathe". I'm still recovering to the point of being able to set goals for a job search (or other plan for the future).

I also believe that the expectation that everyone solo, all the time, led to technical and product failures at multiple companies.

There’s a response I often get at this point, especially from people who were managers in the organization at the time:

“But Jim, teams weren’t expected to solo all the time. You might have been assigned tasks, but you were given leeway to accomplish those how you wanted, and management didn’t require people to solo all the time. If a team wanted to solo less, they could.”

This is true. Engineers and teams had a lot more freedom than they realized they had. I spent a lot of my time there helping people realize that. I would often spend one or two days a week pairing for at least part of the day.

And yet.

Engineers had individual laptops and sometimes set up them according to individual preferences.

We were Tech Company Engineers, and one of the things that made us Tech Company Engineers was that we soloed.

One of the great and terrible things about Tech, is that it operationalizes peer pressure. It harnesses drives that humans have, drives for identity and belonging, in the service of producing software. These forces were only tenuously under management control.

So I soloed all day, most days, for about five years. This had a lot of upsides, far more than I can list here. Soloing really develops being able to plan out and execute a task, giving yourself time to think through a problem, understanding the technologies you are using, and developing proficiency which might not happen if you are leaning on your pair (sometimes more than you realize). The impact of soloing, especially soloing that much, goes much deeper than its impact on the code, on the particular work the team delivers that week.

(Here would go an anecdote about how people who solo have self confidence and mastery.... sorry I'm not thinking of an immediate analogue to the Overcooked example in the original article).

We take the time to understand a problem so we can make informed decisions.

This is the real power of soloing, intense soloing. Understanding what you’re doing, and adjusting it based on further research or the results of experiments, becomes automatic. For someone who thrived on pairing and loved it when there was a strong team spirit, this was an almost psychedelic experience. I transcended the limitations of the people I was working with and discovered that I was able to accomplish things.

I remember once, looking over a large office filled with people working at workstations, and thinking, “This is the most talented collection of people I have ever seen in one space.” I understand why engineers so fiercely defend their right to hack.

But: cognitive impairment.

Months where I struggled to meet my own basic needs.

Soloing requires putting up a facade of self-sufficiency, to management and the rest of the organization, for hours at a time. Being able to manage oneself, both physically and mentally. I had to manage my space, my decisions, my thought processes, and often my feelings on my own.

This never stopped being draining. Even with an easy team, where I had clear goals and could accomplish things without effort, soloing well requires staying engaged with my environment, with what I'm supposed to be doing. No retreating into sensory experiences, no checking my phone, no wandering off or getting distracted. Maintaining that level of focus for hours at at time was thrilling, but it also required a serious exercise of will.

There were some teams where it required more than will. I had to fight to stay engaged. I had to develop skills. There were people with whom I disagreed, but with whom I struggled to resolve those agreements. There were people who didn’t put nearly as much thought into my experience as I was putting into theirs. There were people who expected me to “just make it work”, despite an ill-defined and not-yet implemented interface that I was expected to use. There were people who just made me anxious or uncomfortable.

I had to confront a lot of my fears about myself. I had to learn to show someone else what I could and could not do, my limitations as a human and a software engineer.

Over time, over years, soloing wore me down. Took a little bit more each day than I could recover. Until my life was working, and recovering from work, and then working some more.

And then the pandemic happened. Overnight, suddenly, I was performing this daily act of will without the support of the office, without going out to lunch at my favorite restaurant, without anyone to talk to except in scheduled meetings, while the world burned down around me.

I crumpled. I stopped being able to solo. Stopped being able to have a conversation.

This wasn’t everyone’s experience with soloing. In a two-by-two grid where one axis is “sensitive to the demands of soloing” and the other is “time committed to soloing” I’m hugging the upper left hand corner. My experience was extreme.

And yet.

Tech companies, I’m told, have reputations as “burnout factories.” Many people left my companies, at least a few of them to escape from the demands of daily soloing. People who love soloing, who see the benefit of it, but who despite all those benefits are tired.

There are people who can solo indefinitely, for years. Who don’t experience the most demanding version of it often. Whose recovery capacity comfortably outpaces the demand. A lot of those people, at most tech companies, end up in management roles, in leadership roles, and then they miss coding. Many places I've worked have had a leadership staff that, even when they believed me when about how demanding soloing was for me, couldn’t really see it themselves. Some of them even tried to find ways to enable me to solo less, but I'm not sure they understood all the forces which were making it hard to do anything but solo.

We’ve all heard the bad reasons not to solo. "You are just cowboy coding." "You can't think about anyone other than yourself and your pet project." "You don't care about doing things well."

I’ve dismissed people making those arguments as fundamentally dogmatic, unwilling to do the hard work of real software development.

Now, though, I hear those objections, and I hear fear. A fear that I share. A fear of exposing my vulnerability, my ignorance, my soft parts, and a fear of the cost of that exposure, of the cost to my mind and my body of subjecting myself to that exposure, day after day, in exchange for a paycheck.

Underneath the urge to dismiss those concerns, I hear another fear. A fear that soloing is too hard, that people wouldn’t choose to do it if they weren’t corralled into it by individual goals and hiring and promotion processes which emphasize "yes, but what did you do personally?". That a “soloing culture” is such a delicate wisp of a thing that if you allowed engineers to solve problems together, they would abandon soloing immediately.

What did we miss out on, by failing to make more space for people not to solo? By treating this soloing culture as something so fragile, and so precious?