I paid for yesterday’s light-meeting day with a heavy-meeting day today. Well-played, calendar gods. But I also read some great content, and even had time for some quick demos about AI-generated data insights and attached volumes on a serverless app.
[blog] Gemini Chat App. Simon used Claude to write a small app that uses our latest Gemini 1.5 model versions. He also opines that people who don’t see value in using AI assistance for programming are missing something.
[article] Why Cynics Are Less Likely to Succeed. It’s not hard to be cynical, but operating in a mode of trust and cooperation is not just good for your mental well-being, it’s better for your career.
[blog] What is the Open Source Alternative to CockroachDB? When license changes happen, other vendors/projects jump in. Denis at Yugabyte offers up a case for using their database as a drop-in replacement.
[blog] Managing Angular. This is a high level view from the product lead for the popular JavaScript framework. OSS management is quite the job, whether you’re solo or working inside a big tech firm.
[article] Speak, Code, Deploy: What if voice was your primary tool for coding? I dunno. I’m hoping speech-to-text and chatbots are transient interfaces with AI. At least for the masses that don’t need that for accessibility reasons. I personally don’t want to talk to my computer, or be stuck “chatting” to get my work done.
I had a very light meeting day today, which messed with my head. But, it was great to answer all my email, write a blog post, do some research, and work on upcoming presentations.
[blog] Routines and habit stacking. Tom looks at incorporating goals into current routines, and piggybacking on existing success.
[article] Does Market Share Still Matter? Do “market leaders” have the most efficiency, market power, and quality? Or do highly digital firms have similar profitability to the market leaders? Interesting research.
[article] Profitable on day one! What does it even mean to be “profitable? Jason encourages us to use the term correctly.
[blog] A Year of Project IDX. If you haven’t checked this out, at least give it a scan. IDX is an interesting developer environment, and I’ve used it to build a few apps.
[blog] Friction Logs. This is about the process of using products, recording the experience and papercuts that come with it, and sending that feedback to those who fix it.
[article] Why AI can’t spell ‘strawberry’. Such an interesting problem! I tried this scenario with the latest Gemini Flash models we released today, and it did indeed answer correctly.
What advice do you get if you’re lugging around a lot of financial debt? Many folks will tell you to start purging expenses. Stop eating out at restaurants, go down to one family car, cancel streaming subscriptions, and sell unnecessary luxuries. For some reason, I don’t see the same aggressive advice when it comes to technical debt. I hear soft language around “optimization” or “management” versus assertive stances that take a meat cleaver to your architectural excesses.
What is architectural debt? I’m thinking about bloated software portfolios where you’re carrying eight products in every category. Brittle automation that only partially works and still requires manual workarounds and black magic. Unique customizations to packaged software that’s now keeping you from being able to upgrade to modern versions. Also half-finished “ivory tower” designs where the complex distributed system isn’t fully in place, and may never be. You might have too much coupling, too little coupling, unsupported frameworks, and all sorts of things that make deployments slow, maintenance expensive, and wholesale improvements impossible.
This stuff matters. The latest StackOverflow developer survey shows that the most common frustration is the “amount of technical debt.” It’s wasting up to eight hours a week for each developer! Number two and three are around stack complexity. Your code and architectural tech debt is slowing down your release velocity, creating attrition with your best employees, and limiting how much you can invest in new tech areas. It’s well-past time to simplify by purging architecture components that have built up (and calcified) over time. Let’s write bigger checks to pay down this debt faster.
Explore these four areas, all focused on simplification. There are obviously tradeoffs and cost with each suggestion, but you’re not going to make meaningful progress by being timid. Note there are other dimensions to fixing tech debt besides simplification, but that’s one I see discussed the least often. I’ll use Google Cloud to offer some examples of how you might specifically tackle each, given we’re the best cloud for those making a firm shift away from legacy tech debt.
1. Stop moving so much data around.
If you zoom out on your architecture, how many components do you have that get data from point A to point B? I’d bet that you have lots of ETL pipelines to consolidate data into a warehouse or data lake, messaging and event processing solutions to shunt data around, and even API calls that suck data from one system into another. That’s a lot of machinery you have to create, update, and manage every day.
Can you get rid of some of this? Can you access more of the data where it rests, versus copying it all over the place? Or use software that act on data in different ways without forcing you to migrate it for further processing? I think so.
Let’s see some examples.
Perform analytical queries against data sitting in different places? Google Cloud supports that with BigQuery Omni. We run BigQuery in AWS and Azure so that you can access data at rest, and not be forced to consolidate it in a single data lake. Here, I have an Excel file sitting in an Azure blob storage account. I could copy that data over to Google Cloud, but that’s more components for me to create and manage.
Rather, I can set up a pointer to Azure from within BigQuery, and treat it like any other table. The data is processed in Azure, and only summary info travels across the wire.
You might say “that’s cool, but I have related data in another cloud, so I’d have to move it anyway to do joins and such.” You’d think so. But we also offer cross-cloud joins with BigQuery Omni. Check this out. I’ve got that employee data in Azure, but timesheet data in Google Cloud.
With a single SQL statement, I’m joining data across clouds. No data movement required. Less debt.
Enrich data in analytical queries from outside databases? You might have ETL jobs in place to bring reference data into your data warehouse to supplement what’s already there. That may be unnecessary.
With BigQuery’s Federated Queries, I can reach live into PostgreSQL, MySQL, Cloud Spanner, and even SAP Datasphere sources. Access data where it rests. Here, I’m using the EXTERNAL_QUERY function to retrieve data from a Cloud SQL database instance.
I could use that syntax to perform joins, and do all sorts of things without ever moving data around.
Perform complex SQL analytics against log data? Does your architecture have data copying jobs for operational data? Maybe to get it into a system where you can perform SQL queries against logs? There’s a better way.
You can’t avoid moving data around. It’s often required. But I’m fairly sure that through smart product selection and some redesign of the architecture, you could eliminate a lot of unnecessary traffic.
2. Compress the stack by removing duplicative components.
Break out the chainsaw. Do you have multiple products for each software category? Or too many fine-grained categories full of best-of-breed technology? It’s time to trim.
My former colleague Josh McKenty used to say something along the lines of “if it’s emerging buy a few, it’s a mature, no more than two.”
You don’t need a dozen project management software products. Or more than two relational database platforms. In many cases, you can use multi-purpose services and embrace “good enough.”
There should be a fifteen day cooling off period before you buy a specialized vector database. Just use PostgreSQL. Or, any number of existing databases that now support vector capabilities. Maybe you can even skip RAG-based solutions (and infrastructure) all together for certain use cases and just use Gemini with its long context.
You could use Spanner Graph instead of a dedicated graph database, or Artifact Registry as a single place for OS and application packages.
I’m keen on the new continuous queries for BigQuery where you can do stream analytics and processing as data comes into the warehouse. Enrich data, call AI models, and more. Instead of a separate service or component, it’s just part of the BigQuery engine. Turn off some stuff?
I suspect that this one is among the hardest for folks to act upon. We often hold onto technology because it’s familiar, or even because of misplaced loyalty. But be bold. Simplify your stack by getting rid of technology that’s no longer differentiated. Make a goal of having 30% fewer software products or platforms in your architecture in 2025.
3. Replace hyper-customized software and automation with managed services and vanilla infrastructure.
Hear me out. You’re not that unique. There are a handful of things that your company does which are the “secret sauce” for your success, and the rest is the same as everyone else.
More often than not, you should be fitting your team to the software, not your software to the team. I’ve personally configured and extended packaged software to a point that it was unrecognizable. For what? Because we thought our customer service intake process was SO MUCH different than anyone else’s? It wasn’t. So much tech debt happens because we want to shape technology to our existing requirements, or we want to avoid “lock-in” by committing to a vendor’s way of doing things. I think both are misguided.
I read a lot of annual reports from public companies. I’ve never seen “we slayed at Kubernetes this year” called out. Nobody cares. A cleverly scripted, hyper-customized setup that looks like the CNCF landscape diagram is more boat anchor than accelerator. Consider switching a fully automated managed cluster in something like GKE Autopilot. Pay per pod, and get automatic upgrades, secure-by-default configurations, and a host of GKE Enterprise features to create sameness across clusters.
Or thank-and-retire that customized or legacy workflow engine (code framework, or software product) that only four people actually understand. Use a nicely API-enabled managed product with useful control-flow actions, or a full-fledged cloud-hosted integration engine.
You probably don’t need a customized database, caching solution, or even CI/CD stack. These are all super mature solution spaces, where whatever is provided out of the box is likely suitable for what you really need.
4. Tone it down on the microservices and distributed systems.
Look, I get excited about technology and want to use all the latest things. But it’s often overkill, especially in the early (or late) stages of a product.
You simply don’t need a couple dozen serverless functions to serve a static web app. Simmer down. Or a big complex JavaScript framework when your site has a pair of pages. So much technical debt comes from over-engineering systems to use the latest patterns and technology, when the classic ones will do.
Smash most of your serverless functions back into an “app” hosted in Cloud Run. Fewer moving parts, and all the agility you want. Use vanilla JavaScript where you can. Use small, geo-located databases until you MUST to do cross-region or global replication. Don’t build “developer platforms” and IDPs until you actually need them.
I’m not going all DHH on you, but most folks would be better off defaulting to more monolithic systems running on a server or two. We’ve all over-distributed too many services and created unnecessarily complex architectures that are now brittle or impossible to understand. If you need the scale and resilience of distributed systems RIGHT NOW then go build one. But most of us have gotten burned from premature optimization because we assumed that our system had to handle 100x user growth overnight.
Wrap Up
Every company has tech debt, whether the business is 100 years old or started last week. Google has it, big banks have it, the governments have it, and YC companies have it. And “managing it” is probably a responsible thing to do. But sometimes, when you need to make a step-function improvement in how you work, incremental changes aren’t good enough. Simplify by removing the cruft, and take big cuts out of your architecture to do it!
I respected summer this weekend by being outside and rejecting any retail efforts to make me buy Halloween decorations or drink pumpkin spice drinks. Stay strong!
[article] Strangler Fig. I’ll be honest with you. I’ve referred to this design pattern a few times, without having any idea why it was called that. Now I do, in this updated post from Martin.
[blog] Gemini Bounding Box Visualization. Simon uses Claude to write some code against Gemini APIs that produce a bounding box around specified objects. Cool!
[paper] Measuring Developer Goals. This new paper highlights 30 developers goals that we at Google track to help improve the developer experience for our engineers.
[blog] Manager Antipatterns. Ted calls out things that most of us in management roles have done (or are doing), and should actively try and rectify.
Walked in the door at midnight last night after flying home from a business trip, and that 5am alarm was rough! But today was a productive catch-up day, and I hope many of you are enjoying the tail end of Summer.
[blog] Structured logging in Spring Boot 3.4. A well-instrumented system is better than one that isn’t. This post shows how Java apps using Spring Boot can have nicely structured logs with little effort.
[blog] Gemma explained: What’s new in Gemma 2. The architecture for this latest open model version differs from its predecessor. This post explains the differences and the impact they had.
[article] Rackspace Goes All In – Again – On OpenStack. I can’t say I’ve heard much about either lately. But private cloud fans may like seeing a fully managed OpenStack offering form Rackspace.
I had an enjoyable day here in Kansas City, meeting with four different customers. It was very helpful to get reminders of “real world” challenges and use cases. Flying home now, and back in the San Diego saddle tomorrow.
[blog] Cloud Run GPU: Make your LLMs serverless. I’m starting to think that this is a big deal. Easier access to LLM hosting and experimentation will only expand the realm of use cases.
[blog] Serverless Sucks? But why all this serverless stuff? Just use servers and make life easier. Not so fast, says Derek.
[blog] gRPC Communication Between Go and Python. If all your app components are written in the same language, kudos. That makes life easier. But for most everyone else, we need to think about interop and how to get these pieces talking. Here’s a good example.
I made it to Kansas, and I’ve got a bunch of fun demos lined up for customers tomorrow. How do you feel about doing (or watching) live demos? I know there’s a risk, but it feels more genuine than playing a recording.
[guide] Use generative AI for utilization management. This reference architecture is targeted at health insurance companies, but it offers a blueprint for anyone doing automation of document processing.
I’ve got a short trip up to Kansas City tomorrow to chat with some customers and do some live demos. I like working at a place that requires me to constantly update my demos to show off new functionality!
[blog] Does Open Source Work as a Long-Term Business Model? It doesn’t look like Yugabyte is going to copy CockroachDB’s license change. This post looks at the rise of PostgreSQL and the risk of positioning against it.
I had a fun weekend with visiting family, and now I’m back to regular work. Today’s reading list has a lot of AI-related content, but coming at it from a variety of angles.
[blog] Is GenAI Falling Short? Not In The Contact Center. I like these two examples. Both automated call summarization and support for “infrequently asked questions” are valid use cases for AI today.
[blog] PyTorch is dead. Long live JAX. Edgy title! But also a very well-reasoned argument for why DeepMind’s JAX framework may be a better long term choice.
[blog] Gemma for Streaming ML with Dataflow. This, however, is a strong use case for AI. Add it to data processing pipelines for actions like sentiment analysis.
[paper] Imagen 3. This looks like a huge jump in performance for text-to-image generation. Read this new research paper for information about this Google model, and some human ratings of generated images.
[blog] Improving Design Reviews at Google. We’re constantly studying data about how we work at Google, and experimenting with ways to make it better. This post summarizes a paper about our efforts to improve the design review process.
[article] The Product Model in Outsourcing. Fascinating piece from Marty and Josh that offers specific advice for those that want more product-oriented thinking from their outsourced/service partners. It requires major changes on both sides!
[article] MIT researchers release a repository of AI risks. This is useful for researchers, customers, and even vendors. I’m sending this to my docs team to see if we should improve our documentation.
[blog] Bash to Terraform with Gemini 1.5 Flash. All those scripts you have laying around might be turning into tech debt. Maybe converting to something like Terraform is more maintainable? Either way, this is an interesting look at how to do that.