I solve real problems at scale and the engineering practices I see on a daily basis are a clown show.
There's little to no basic understanding of networking, distributed systems, databases, etc. 99% of our engineers were hired from their college internships and never worked anywhere else. Industry hires to improve systems rarely last more than a year and it is almost never their fault.
We're in the next tier down from the biggest tech companies and what we do is hardly uncommon among our peers.
I should be shocked that 99% of engineers I deal with treat all resources as infinite bandwidth, 100% uptime, but I'm not. They NIH super hard and write tons of code for things that a docker container running nginx (or similar) would solve in 5 minutes. There's almost no useful testing and worse documentation.
I had to explain to a “senior” engineer the other day (read: a few years experience) why locating a database client in a different geo region from the server is a bad idea (especially when that client is using an ORM that likes to make lots of little calls to the server.)
I also had to argue for changing a system that was reading about 100k small files from cloud storage to use a single compressed file. There seemed to be no awareness that copying 100k files might be inefficient.
One company I was at sold its product claiming any changes you made were “near instantly published” globally. They tried to demo it as such.
The way the engineers built the update/publish operation was synchronous from their primary data center to a number of globally distributed data centers. Publish didn’t “complete” until a receiver in each data center responded with an ACK after parsing and uploading to a nearby region cloud bucket. Any failure/timeout caused the entire transaction across all data centers to retry. All of traffic was over multiple VPNs, hub and spoke style. They built this system in 2020.
They constantly complained and generated incident reports about p95/p99 latencies to the Asia regions. Latencies that were perfectly reasonable when you considered the multiple global round trips that were being made, the size/volume of objects in the publish, set of operations and speed of light.
They swore that because the client UX to publish the change to the primary data center used JavaScript async that the entire process was async. They denied repeatedly that their “all receivers ack complete to succeed” business logic was synchronous. I shit you not.