Productive Signaling: Competitive Software Development, Not Competitive Programming

This post is crossposted from my Substack, Structure and Guarantees, where I explore how formal verification and related ideas might scale to more complex intelligent systems. Here I argue for an opportunity from the rise of AI coding tools: we can replace elite programming competitions, where students write throwaway programs, with competitions instead about creating useful new software, scored on adoption or other measures of economic value created. Then we get strictly more benefit than merely identifying the strong students and training them in programming.

My last article argued that we have an opportunity to enter a rapid cycle of reimplementation of important software, as AI and automation tools broadly lower the cost of development from sufficiently clear high-level specifications. One reason to “rewrite everything” is advances in cybersecurity techniques and knowledge, on the sides of both offense and defense. Another is evolution of the software-hardware interface, requiring substantial rewriting of old code in new languages to stay performance-competitive. There’s just one problem (well, OK, there are several): the software engineers are generally already busy maintaining their existing systems. How do we get this work done, in this era when we still can’t turn over the entire process to automation?

My answer is an instance of a pattern I plan to return to occasionally. Let’s think of the global economy as a distributed system solving an optimization problem to organize the world properly. One of its tools is signaling, where people produce costly, hard-to-fake displays of competence, which are used to evaluate them and partly decide where to slot them into the economy. Clearly that evaluation process produces massive downstream value. However, it also produces some pure “heat” with resources expended on signaling without directly moving the primary metrics of interest to the economy. In biology, signals probably evolve in the first place thanks to their connections to fundamental survival and reproductive success, even if runaway selection can take them far from that territory over time.

With our present-day abstract-thinking capabilities, we can choose to realign signaling deliberately so that it not only provides competence information but also generates first-order economic value. I’ll use the productive signaling header for articles that suggest instances of this pattern. Today we’re considering how to help spur widespread regeneration of important software.

Among top students worldwide in high school and a bit younger, there is a tradition of participation in olympiads like the International Olympiad in Informatics (IOI) or International Mathematics Olympiad (IMO). The basic form is hierarchical competition in particular STEM areas, with more localized competitions choosing winners to advance to competitions with broader geographic scope. For instance, the IOI challenges participants to write computer programs quickly that implement provided requirements, generally focusing on problems framed in terms of simple discrete mathematical objects, with straightforward specifications but nontrivial solutions. (It’s the subject matter that dominates classes commonly called just “algorithms,” but the real scope of different kinds of tricky algorithms worth writing is much broader than we see in classic competitive programming.)

These kinds of competitions are very useful for identifying stars in the respective areas of knowledge. My own institution MIT seems to lean heavily on olympiad records to make decisions in undergraduate admissions, especially for international students. That is, olympiads produce a simple quantitative signal (based on standing in international contests) that is highly legible even to nonexperts, like people in charge of undergraduate admissions. I believe the signal being measured is also very relevant to potential to succeed in related careers. And olympiads are not just a measurement process: training for them helps upskill students and genuinely prepare them to contribute to important social problems in not too many years. We should also acknowledge that, given the centrality of signaling in our own evolution, it’s extremely common for people to hunger for chances to compete successfully and show off their talents, and we should be careful about letting video games become the default outlet!

Call me greedy, but, even given those benefits, I think it would still be fabulous if this experience also helped solve social problems directly. We expect to see approximately a few hundred competitors at a typical olympiad world-finals event, and they are all solving the same problems. The duplication is even greater for the sum of all more-localized competitions that determine competitors at world finals. We only need one competent team to solve a problem to realize the full social value of generated code for that problem! So the first big problem with this structure is that, from the standpoint of producing useful code, it creates massively redundant work, where we can subtract essentially all teams/competitors but the most-successful for a given problem and achieve the same practical outcome. We can see why this uniformity evolved, as it makes judging tractable, but hold that thought.

The other big problem is that solutions to the problems assigned in olympiads rarely create immediate social value, independently of signals about competitor competence levels. Pretty much by definition, these problems were already solved when the competition started! The organizers would be embarrassed to assign an unsolvable problem, and that pitfall is hard to avoid without solving in advance. There is also the scale of the problems that makes them unlikely to have big social impact. The time limitations of these contests push toward limited scope, considering how important real-world problems often occupy teams of software engineers for years, even before first wide releases. AI coding assistance is changing that equation, but please also hold that thought!

My alternative proposal is a competition structure around creating useful new open-source software. Developing a new and better program in an existing category is explicitly encouraged, and in fact it connects to our initial motivation to kick off widespread regeneration of important software, drawing on new wisdom in cybersecurity and elsewhere. Olympiads produce many variants of the same program, but they are all promptly thrown away, while I’m suggesting a contest that develops variants with increasing real-world value. If we can just judge competitors properly and summarize results in a form highly legible to nonexperts who make decisions like undergraduate admissions, then we are in business: we have harnessed signaling in a way that produces a socially valuable first-order output, not just the handy signal. (As a quick parenthetical tangent, AI is changing the educational landscape so much that it isn’t clear the admissions process we know today will last much longer, but if it goes away and isn’t replaced by other similar talent evaluations, humans have probably stopped being competitive to do knowledge work. Oops!)

Comparison.png

I should add that, through showing a draft of this article to an LLM, I learned for the first time about Google Code-in, which brokered between high-school-ish students and open-source projects that provided tasks within their code bases. On the one hand, I frequently interact with students just past this age range applying for research roles, and since I hadn’t heard of Code-in before, it is unlikely to have become a common tool for signaling programming ability. However, the fact that it continued for 10 years, with almost 15,000 students participating, provides some directional evidence that this sort of initiative can scale. Like the Google Summer of Code, which brings students into established open-source projects for quasi-internships, Code-in focused on projects that had been around for a while. Code-in wound down in January 2020 (did they know the pandemic was coming?!), just a little too soon to benefit from the new wave of AI-coding tools. It is already very feasible for the best-qualified students to create their own useful new software packages from scratch, and we should take advantage of that more-granular signal. Maybe it is only in-reach each year for approximately the number of students who participate in olympiad world championships or even earn top distinction in them, but that’s still an important population. Also, while that contributor count may not represent a big-enough workforce for e.g. a wide retrofitting of important software with new security protections, we can hope that advances in AI coding tools grow the population with time, and the contest structure can evolve to take advantage. (In this case, there is a helpful correlation, where when the risks from AI models finding vulnerabilities grow, the tools probably also improve to allow larger contests to maintain quality standards.)

Which details need to be nailed down to make this kind of competition work?

Open-source maintainers are already suffering from firehoses of AI-generated slop “contributions,” and you want to add to the problem? Don’t worry: this proposal is for creation of new software projects and doesn’t require interfering with the normal workings of established ones!

We are talking about creating programs worthy of mass adoption. Do we really want to entrust that task to high-school students? I’m not talking about a standardized test that all high-school students need to take. A very small fraction of high-school-equivalent students worldwide participate in the feeder programs for IOI. It only needs to be feasible for that echelon of students to build programs that at least make good starting points for broader open-source contributions, perhaps overwhelmingly by professionals, after improved fit for an important problem is demonstrated. And they get to use all the latest AI-powered software-engineering tools! We are still getting used to the power and engineering implications for those tools, but I don’t think it’s crazy to imagine small student teams or even solo competitors making really valuable and solid stuff over manageable time periods (which can nonetheless be at least months long, in stark contrast to durations of olympiad contests). Also, naturally, all of these efforts should be associated with formal verification that permits checking of quality for resulting code, but let me not get too sidetracked by that hobby horse.

So we have each team writing its own new program, covering a wide range of domains to maximize social value on success. How on earth do we judge the results effectively? Just getting AI to do the judging is probably a helpful start but also vulnerable to pernicious gaming of the system, so we are going to need human judges. (Again, maybe we don’t need human judges, but then we’ve probably crossed over to the regime where, for better or worse, we don’t need to train humans in programming or evaluate their skills in it.) We can take advantage of the same hierarchy already found in olympiads: a succession of increasingly global competitions, where wider-scope competitions are fair to assume as including competitors with higher average competence. The earliest stages could even keep the conventional programming-contest format, switching to developing novel programs later on, when we don’t need as many judges. Competitors need to earn the reputations that their programs are worth evaluating. However, also think of companies eager to meet the best potential hires early. They may volunteer their engineers’ judging time, even relatively early in the hierarchy of competitions. It may even be possible to crowdsource evaluation massively, using a metric like GitHub stars to judge outcomes, though the chances to game that metric (e.g. with sock puppets) are clear. A proxy for real economic value derived by users would be ideal.

You’re talking about evaluating the program that is produced, but we know that it’s actually AI writing most of the code these days, and the extent of that phenomenon will only increase. How do we know judging isn’t really just evaluating how much token budget a kid’s parents endowed financially? If we get fully automated generation of valuable new open-source software, that’s also a great outcome! Then we probably don’t need mass training and evaluation for programming skills anymore. If an important role remains for human guidance, then capacity in that role is precisely what we want to measure and what a judging process here should be able to get at. That still leaves what seem to me the biggest challenges to resolve, including the difficulty of effective expert review of reams of AI-generated code, to assign scores. Maybe this domain could be a good testing ground for techniques that will matter even more for industrial software-engineering teams? We can also hope costs of AI coding assistance are headed downward (though I have my suspicions that this goal will require moving away from deep learning, which is inherently slow and expensive).

OK, but, seriously, what are the human judges actually doing to determine what skills the contestants have demonstrated? I agree it’s a tough question! We can see clearly the immense payoff from traditional programming competitions giving everyone the same problems and grading based on test cases passed, though the cost is not just lack of production of adoption-worthy code but also practice only on programming tasks that are unrepresentative of most real development in important ways. For now (before neural implants, say!), it should be safe to have face-to-face interviews with students, going through their code and asking the right questions about it. It’s a labor- and time-intensive process, which is why it’s important that relatively few contestants reach this level of the competition. We would also need some kind of rubric, a scoring formula that helps address the massive heterogeneity in what creates value across different kinds of programs, and this rubric also sounds nontrivial to create. In general, I think the saving grace is that this scoring aspect doesn’t need to be very precise; it just needs to identify a cohort of participants who have demonstrated enough competence to be worth inviting into the next-higher level of participation (e.g. university admission or a first job). Also, it should be possible to borrow good ideas from science fairs, which also involve judging within a pretty open-ended design space of projects selected by students and also are already recognized as a somewhat common signaling currency by university admissions.

Are there even enough high-school staff out there who are qualified to help students succeed in this kind of competition? Maybe not, but there’s a decent chance that generative AI provides a good substitute! Students can go beyond just asking AI to write code by also asking AI questions that prepare students to do more work on their own.

How will students find which problems are worth solving with software, anyway? There could be some shared online resource with proposal of problems and voting on which ones matter, by suitably vetted audiences. Companies could propose problems alongside offers to judge solutions, perhaps augmented by something like automated test cases. However, especially with the way things are heading with AI coding assistance, choosing the right problems makes up an increasingly large fraction of effective software development. We may want to measure exactly this skill, even if it is very rare in the student population, even if research to find the right problem comes to dominate time spent on the competition!

How will students know they should spend time on this weird new thing? It’ll take some time to bootstrap marketing, I’m sure. Our brains evolved to expect social contexts like those in hunter-gatherer bands, where sets of skills worth learning stayed remarkably constant (from a present-day perspective) even over thousands of years. Clear guidelines and developmental rituals were provided, to guide young people through skill development. Olympiads have tied into this cognitive machinery today: students look around for what people do to develop skills and get credit for developing them, and they head for the competitions suited to their talents. Schools build up resources to help students and encourage them into the proper paths, motivated partly by the school community’s chance to get some signaling credit for elevating champions. If there are enough/important evaluators out there ready to use this signal, the rest should follow over time.

I sneakily saved a disclaimer for the end. I personally stuck with competition programming only minimally in my high-school days. Instead, I spent my time writing programs and getting people online to use them. Readers should be suspicious that I’m arguing for a change in signaling regime that disproportionately benefits people like me. I hope there’s still enough of a case that the social benefit is large!

It’s interesting to speculate (but I’m less qualified to do it) on olympiads in other domains. An especially intriguing one for me is math. Do we want students all going after different open problems to prove with Lean? Is it even conceivable that this early-career cohort could exhibit collective productivity comparable to what the whole international math community was pulling off five years ago?

There are plenty of logistical challenges left to sort out to implement these ideas; my goal is just to get the conversation started. My next post will return to the idea of frequent, rapid, automation-supported redevelopment of software, trying to fill in some of the holes pointed out by the great reader comments on the previous article (ironically, that link won’t take you to them, but they’re on LinkedIn, GeekNews, LessWrong, and ACX).

AI Article