It has more than 2M lines of C code. Almost all libraries were written from scratch.
Every professional security researcher reading this just raised an eyebrow and thought to themselves, "That's a vulnerable application." In fact, your team (or a similar one at AOL) did introduce vulnerabilities[1][2]. If we assume that, as you imply, AOL really was exclusively composed of above average C developers writing high quality code, this should really just demonstrate the point for us.
I'm sure your team wrote high quality C code (or at the very least, tried very hard to write high quality C code), and as someone who likes C I also dislike the meme that C should simply not be written, no exceptions. I'd probably change it to, "Don't write C unless you need extreme performance, have great engineering resources and you'll generate unmitigated financial returns such to make any reasonable business risk assessment moot." That means virtually the same thing for anyone reading it, but it sounds a lot less sexy and memorable as far as programming aphorisms go.
I have never professionally audited a C project and not found a vulnerability. Most of my friends in this industry would likely say the same. I've never seen a serious audit of a "large" (100kloc or more) C project by a serious, reputable firm with no vulnerabilities reported. If you wrote 2 million lines of C code, if you wrote libraries from scratch, you have serious vulnerabilities. Otherwise, your engineering practices and resources exceed NASA's in formal verification and correctness.
I believe it is theoretically possible to write 100% safe C code, for some approximation of that ideal that leaves aside the ivory tower definition ("the only safe machine is the one which is off"). But I also believe this is so expensive and resource intensive (secure coding standards, audits, attempts at formal verification and provable correctness) that it's just generally not realistic in practice.
The reason I am saying all of this is not to attack you or your sense of worth as a C developer. Rather, I'd like you to consider that you simply have no idea how many vulnerabilities exist in ICQ. In fact, I found eight different security reports for ICQ, linked at the bottom. One of the more serious ones allowed arbitrary email retrieval, which I'm willing to bet was introduced because your team liked to develop in-house libraries and decided to do that for POP3.
C is a language which is both very powerful and which requires constant vigilance while coding. It is a language which makes pushing a buffer overflow to production on a Friday at 3 pm very easy. I bet your team was comparatively well-versed in C vulnerabilities for the time, but the very fact that you rolled libraries from scratch tells me you already have a high probability of errors showing up. Rolling a library/framework in-house is basically the first thing security researchers look for in an audit, because the third party ones are (usually) more secure due to their exposure. When you further consider the probability of third party software you did use becoming vulnerable in the future, changing compiler optimizations introducing new vulnerabilities and the sheer size of the SEI CERT Secure Coding Standard itself, you are left with a very dangerous risk profile.
We cannot expect C to be a safe language for general use if you need to have the rigor of one of the best research organizations in the world coupled with the top 0.01% of all C programmers. It just isn't feasible.
> Every professional security researcher reading this just raised an eyebrow and thought to themselves, "That's a vulnerable application."
I think the reason why every seasoned security researcher I've met also happens to be a heavy drinker or is damaged in some way is because they're in a war of attrition. You want perfect security but it will never happen. These are ultimately mutating, register-based machines. We will probably never be certain that with a sufficient level abstraction a program can be written which will never execute an invalid instruction or be manipulated to reveal hidden information.
Where the theory hits the road is where the action happens.
Which is how we end up with this wide spectrum of acceptable tolerances to security. Holistic verification of systems is extremely costly but necessary where human lives matter. However if someone finds a weird side-channel attack in a bitmap parsing library I think we can be more forgiving.
The whole idea that C programs are insecure by default and can never be secure is where theory wants to ignore the harsh realities. We can write languages with tighter constraints on the verification of the programs they create which will lower the risk of most security exploits by huge margins... but we have to trade something away for the benefit. The immediate costs being run-time performance or qualitative things like maintainability.
What I ultimately think will make these poor security researchers feel better is liability. Having a real system and standard in place for professional practices will at least let us soak up the damage will force us to consider security and scrutinize our code.
> The immediate costs being run-time performance or qualitative things like maintainability.
I don't think that's true. There's no reason why memory safety has to cost either performance or maintainability.
It feels like there's some sort of fundamental dichotomy because we didn't know how to do it in 1980, and we're still using languages from 1980, but our knowledge has advanced since then.
I probably shouldn't get in on this heated discussion ... but C++ aims to be a safer (higher level) C, at no runtime cost. I think it mostly succeeds at this.
I think C++ can only get so far, without deprecating some of C. Or at least C style code, should come with a big warning from the compiler:
"You are using an unsafe language feature. Are you absolutely sure there is no bug here, and there are no way to use safe language features instead?"
In libraries there might be cases where C-style code like manual pointer arithmetic, c arrays, manual memory management, c-style casts, or *void pointers, uninitialized variables, etc are necessary. But in user code most often there are safer replacements.
There are no such provably safe programs. I assume that you are thinking about programs that have been verified by formal proof systems.
They are not proven to be safe, they are proven to adhere to whatever the proof system managed to prove. This shifts the possible vulnerabilities away from the actual code and onto the proof system and the assumptions that the proof system makes.
For example, take a C program that was proven to be absolutely free of buffer overflows. And therefore labelled by the proof system as 'secure'. But, unbeknownst to the proof system, the C program also interprets user-supplied input as a format string! So it's still vulnerable to format string exploits.
Correct proof systems probably add a huge margin of security compared with the current security standards, but it's not absolute.
There are all kinds of people doing security research. In my experience with some of those people that I've met it doesn't seem like the challenges they have to contend with are not unrelated to the stress caused by the work that they do.
Sort of like how someone who works with giant metal stamping presses is likely to be missing a couple of fingers if they've worked long enough with them (and in an environment where safety regulations are too relaxed).
These are also some of the wonderful people in my life and I enjoy them very much.
"Don't write C unless you need extreme performance,
This is the wrong mindset and has been the guiding principle behind over a generation of poor programming discipline.
You might be able to rationalize that you don't always need extreme performance, but it's not about performance, it's about efficiency. Efficiency always matters. We live on a small planet with finite resources. Today, the majority of computing is happening with handheld portable devices with tiny batteries. Every wasted byte and CPU cycle costs power, runtime, and usability.
Even if you're coding for a desktop PC or a server application - every bit of waste costs electricity, generates heat, increases cooling load. Every bit of waste costs the user real money, even if they don't see the time that you wasted for them.
Performance is a side-effect. The only thing that matters is efficiency, and it matters all the time.
Most of today's programming problems come down to laziness and fear of "premature optimization" (which is obviously bunk). It is possible to write perfect C code that runs correctly the first time and every time. You do that by planning ahead. Formal specs are part of this.
The same consideration applies to writing, in general. I witnessed this first hand, when I worked at the computer labs at the University of Michigan back in the 1980s. When all anyone had was an electric typewriter, they planned their papers out well in advance - detailed outlines, with logically designed introductions and conclusions. This was a requirement, because once you got into actually typing, correcting a mistake could be expensive or impossible without starting over from scratch.
When we only had Wordstar available, and users came in with the inevitable questions and requests for help, you would still see well-organized papers written with a clear logical flow.
In 1985 with the advent of the Macintosh and MacWrite, all of that changed: People discovered that they could edit at will, copy and paste and move text around on the fly. And so they did - throwing planning out the window. And the quality of papers we saw at the help desk plummeted. Surveys from those years confirmed it too - paradoxically, the writing scores in the better-funded departments dropped rapidly, commensurate with their rate of introduction of Macs.
Planning ahead, making sure you know what you need to write, before you start writing any actual lines of code, prevents most problems from ever happening. It's a pretty simple discipline, but one that's lost on most people today.
> Most of today's programming problems come down to laziness and fear of "premature optimization" (which is obviously bunk). It is possible to write perfect C code that runs correctly the first time and every time. You do that by planning ahead. Formal specs are part of this.
If "planning ahead" were enough to prevent security vulnerabilities in practice, someone would have done it by now. Instead, we've had C for upwards of 30 years and everyone, "bad programmers", "good programmers" and "10xers" alike, keeps introducing the same vulnerabilities.
It's a nice theory, but we've been trying for over 30 years to make it work, and it keeps failing again and again. I think it's time to admit that C as a secure language (modulo the extremely expensive formal coding practices used in avionics and whatnot) has failed.
Sorry, but your statement makes no sense. That's like saying "we've been trying for over 5000 years to make hammers work, and people keep smashing their thumbs and fingers again and again. I think it's time to admit that hammers as a safe tool have failed."
Security is not a property of the tool, it is a property of the tool's wielder.
I'm not talking about NPE (though I would be happy to elaborate on how not having null pointers also reduces crashes in the wild). I'm talking about exploitable and/or non-type-safe memory safety problems.
Writing in Java is an effective way to reduce remote code execution vulnerabilities over writing in C. Statistics have shown this over and over again.
The very fact that NPE exists, in a language which doesn't have pointers, running on a JVM that doesn't have pointers, is evidence that the entire model of "doesn't have pointers" is a sham. The reality of computer architecture is that data resides in memory, memory is an array of bytes, and bytes have addresses (i.e., pointers) and denying this fact is impossible. Even in a pure pointerless system like Java. And if there are pointers in the system, somewhere, anywhere, they are exploitable.
Remote code execution vulnerabilities in C are less a feature of the language syntax and mostly due to the abysmally poor C library, which is regrettably part of "C The Language" - but no one forces you to use it - there are better libraries. Also, it's an artifact of implementation - just as Java's supposed immunity is an artifact of implementation, not an inherent feature of the language syntax.
As for RCEs in C code - these would be easily avoided if the implementation used two separate stacks - one for function parameters, and one for return addresses. If user code can only overwrite passive data, stack smashing attacks would be totally ineffective. Again - there's nothing in the C language specification that dictates or defines this particular vulnerability. It is solely an internal implementation decision.
> As for RCEs in C code - these would be easily avoided if the implementation used two separate stacks - one for function parameters, and one for return addresses.
No. Increasingly, RCEs are due to vtable confusion resulting from use after free.
The trouble here is that the efficiency you hold up so highly is marginal in most cases. Making the computer do work to make human lives easier isn't a waste of resources, it's the only reason computers exist in the first place.
Totally agree with your latter statement. But the overall assertion still leads people down misguided paths. Modern OSs like MacOS, Windows, and desktop Linux are presumably intended to make using computers easier. Each new release adds reams of new code implementing new features intended to improve ease of use. Yet the net effect is often negative, because the sheer amount of code added makes the overall system significantly slower.
Meanwhile, new programming languages abound, presumably to make programmers' lives easier. But these developments often completely lose sight of the users. Python is easier to write, but a program written in python uses 1000x the CPU and memory resources of a program written in C. It does the job, but slowly, and consuming so much that the computer is unable to do anything else productive. Meanwhile, it's questionable whether programmers' lives have been improved either, because they spend too much time chasing the next new tool instead of growing expertise in just one.
As for marginal efficiency gains - that is only a symptom of the latter developer problem. E.g., LMDB is pure C, less than 8KLOCs, and is orders of magnitude faster than every other embedded database engine out there. It is also 100% crash-proof and offers a plethora of features that other DB engines lack. All while being able to execute entirely within a CPU's L1 cache. The difference is more than "marginal" - and that difference comes from years of practice with C. You will never gain that kind of experience by being a dilettante and messing with the flavor of the week.
> > It has more than 2M lines of C code. Almost all libraries were written from scratch.
> Every professional security researcher reading this just raised an eyebrow and thought to themselves, "That's a vulnerable application."
That may be true, but substitute "C" with "Python" and wouldn't you have the same reaction? Maybe you would expect it to be slightly less vulnerable, but the key risk that I read in that statement (I am not a security professional) is the "2M" and the "written from scratch". The risk from "C" is secondary.
I would think a 2M line Python application would suffer logic bugs, like the C version. However, I wouldn't worry about indexing off the end of arrays, NULL pointer accesses, etc. which occur in C (sometimes silently) on top of the logic errors.
"I have never professionally audited a C project and not found a vulnerability."
Just out of interest: How many C projects have you audited? And have you ever looked for a relationship between the number of vulnerabilities and the percentage code coverage from automated tests?
Every professional security researcher reading this just raised an eyebrow and thought to themselves, "That's a vulnerable application." In fact, your team (or a similar one at AOL) did introduce vulnerabilities[1][2]. If we assume that, as you imply, AOL really was exclusively composed of above average C developers writing high quality code, this should really just demonstrate the point for us.
I'm sure your team wrote high quality C code (or at the very least, tried very hard to write high quality C code), and as someone who likes C I also dislike the meme that C should simply not be written, no exceptions. I'd probably change it to, "Don't write C unless you need extreme performance, have great engineering resources and you'll generate unmitigated financial returns such to make any reasonable business risk assessment moot." That means virtually the same thing for anyone reading it, but it sounds a lot less sexy and memorable as far as programming aphorisms go.
I have never professionally audited a C project and not found a vulnerability. Most of my friends in this industry would likely say the same. I've never seen a serious audit of a "large" (100kloc or more) C project by a serious, reputable firm with no vulnerabilities reported. If you wrote 2 million lines of C code, if you wrote libraries from scratch, you have serious vulnerabilities. Otherwise, your engineering practices and resources exceed NASA's in formal verification and correctness.
I believe it is theoretically possible to write 100% safe C code, for some approximation of that ideal that leaves aside the ivory tower definition ("the only safe machine is the one which is off"). But I also believe this is so expensive and resource intensive (secure coding standards, audits, attempts at formal verification and provable correctness) that it's just generally not realistic in practice.
The reason I am saying all of this is not to attack you or your sense of worth as a C developer. Rather, I'd like you to consider that you simply have no idea how many vulnerabilities exist in ICQ. In fact, I found eight different security reports for ICQ, linked at the bottom. One of the more serious ones allowed arbitrary email retrieval, which I'm willing to bet was introduced because your team liked to develop in-house libraries and decided to do that for POP3.
C is a language which is both very powerful and which requires constant vigilance while coding. It is a language which makes pushing a buffer overflow to production on a Friday at 3 pm very easy. I bet your team was comparatively well-versed in C vulnerabilities for the time, but the very fact that you rolled libraries from scratch tells me you already have a high probability of errors showing up. Rolling a library/framework in-house is basically the first thing security researchers look for in an audit, because the third party ones are (usually) more secure due to their exposure. When you further consider the probability of third party software you did use becoming vulnerable in the future, changing compiler optimizations introducing new vulnerabilities and the sheer size of the SEI CERT Secure Coding Standard itself, you are left with a very dangerous risk profile.
We cannot expect C to be a safe language for general use if you need to have the rigor of one of the best research organizations in the world coupled with the top 0.01% of all C programmers. It just isn't feasible.
[1]: http://www.coresecurity.com/content/bevy-of-new-icq-vulnerab...
[2]: https://www.cvedetails.com/vulnerability-list/vendor_id-123/...