Remote Labor IndexMeasuring AI Automation of Remote Work

Mantas Mazeika, Alice Gatti, Cristina Menghini, Udari Madhushani Sehwag, Shivam Singhal, Yury Orlovskiy, Steven Basart, Manasi Sharma, Denis Peskoff, Elaine Lau, Jaehyuk Lim, Lachlan Carroll, Alice Blair, Vinaya Sivakumar, Sumana Basu, Brad Kenstler, Yuntao Ma, Julian Michael, Xiaoke Li, Oliver Ingebretsen, Aditya Mehta, Jean Mottola, John Teichmann, Kevin Yu, Zaina Shaik, Adam Khoja, Richard Ren, Jason Hausenloy, Long Phan, Ye Htet, Ankit Aich, Tahseen Rabbani, Vivswan Shah, Andriy Novykov, Felix Binder, Kirill Chugunov, Luis Ramirez, Matias Geralnik, Hernán Mesura, Dean Lee, Ed-Yeremai Hernandez Cardona, Annette Diamond, Summer Yue, Alexandr Wang, Bing Liu, Ernesto Hernandez, Dan HendrycksView original
OverviewExpertalloy voice
Imagine you’re a CTO listening to endless claims that “LLMs are about to automate all white‑collar work.” The question the paper asks is: where’s the empirical, economically grounded evidence? This work, “Remote Labor Index: Measuring AI Automation of Remote Work,” tries to answer that by building a benchmark that is much closer to actual freelance labor markets than anything we have today, and then stress‑testing frontier agents on it. The punchline is stark: according to the authors, the best agent they test is only able to fully automate 2.5% of projects. Let me unpack how they get there, starting from the benchmark design. The Remote Labor Index, or RLI, is a collection of 240 end‑to‑end projects, mostly sourced from Upwork and similar freelance platforms. For each project they capture three artifacts: a brief, a set of input files, and a gold‑standard human deliverable produced by a professional freelancer. Crucially, the freelancer also reports the time spent and the amount they were paid. So each item is not a toy task, but a priced, market‑cleared piece of work. The projects span 23 Upwork subcategories. If you look at the pie chart on page 4, the largest slices are “Other” (31%), Video & Animation (13%), 3D Modeling & CAD (12%), Graphic Design (11%), Audio (10%), Game Development (10%), Architecture (7%), and Product Design (6%). Only a minority lives in the classic “code + text” world where current LLMs shine. The histogram of costs on page 5 shows a log‑wide spread: median cost $200, mean $632.6, maximum $22,500. The corresponding time histogram shows a median of 11.5 hours and mean 28.9 hours, with some projects taking up to 450 hours. Aggregated, RLI represents over 6,000 human‑hours and roughly $144k of labor. The authors emphasize that RLI is significantly harder than prior “economically valuable” benchmarks. The bar chart on page 7 compares mean completion time for three benchmarks—GDPval, HCAST, and RLI—against a sample of real Upwork jobs. RLI virtually overlaps with the Upwork distribution, while the others sit at less than half the mean duration. A second bar chart on that page classifies work into “software engineering,” “research & writing,” and “other.” GDPval and HCAST are dominated by the first two; RLI and Upwork are dominated by “other” – design, multimedia, operations, etc. So the paper argues RLI is a better proxy for the remote labor economy. Let’s turn to the metrics, because they are carefully defined and matter for interpretation. The first metric is automation rate. For a given agent A, you run it on all projects, obtaining AI deliverables {Ai}. For each i, human evaluators compare Ai to the human deliverable Hi and to the brief. They apply a 3‑point scale: 1 means Ai is worse and would not be accepted by a reasonable client; 2 means Ai satisfies the brief as well as Hi and would be accepted; 3 means Ai both satisfies the brief and exceeds Hi in quality. The automation rate is simply Automation(A) = (number of projects with rating 2 or 3) / (total projects). So this is a strict “would the client take it” measure, not a fuzzy similarity score. They also define an Elo‑style relative performance metric. For each pair of agents (Aj, Ak) on a project, evaluators see both AI deliverables plus the human reference and rate two aspects: “project completion” (which is closer to satisfying the brief) and, conditional on both being acceptable, “overall quality.” These are each ternary: Aj better / tie / Ak better. To map this into an Elo system, they use a Bradley–Terry model: each agent m is assigned a utility um, and the probability Aj beats Ak is P(Aj ≻ Ak) = exp(uj) / [exp(uj) + exp(uk)]. From the empirical pairwise preference graph, they fit utilities and then affine‑transform them so the human baseline has Elo 1000 and a 400‑point difference corresponds to 10:1 odds of winning. This makes Elo directly comparable to the automation rate evaluations, since those are just “agent vs human” edges under identical criteria. A third metric is “dollars earned.” This is simply the sum over projects of cost(Hi) for which Ai was accepted. Since the total human bundle is worth $143,991, this gives a clear percentage of the economic mass each agent can capture. Finally, they introduce “autoflation,” an index of cost deflation for the fixed bundle of projects as AI improves. Formally, for project i, let cH(i) be the human cost and {cA(j,i)} be the effective cost for each agent j on that project; if agent j fails, cA(j,i) is set to ∞. The optimal cost to complete the bundle with access to all agents plus humans is C⋆ = Σi min( cH(i), minj cA(j,i) ). The autoflation index is Autoflation = 1 − C⋆ / Σi cH(i). So 0% means no savings relative to humans; 100% would be free automation of all projects. According to the time series on page 17, autoflation is still under 4% as of October 2025. Let’s examine the experimental setup that drives these numbers. They evaluate six state‑of‑the‑art systems: ChatGPT agent, GPT‑5, Claude Sonnet 4.5, Grok 4, Gemini 2.5 Pro, and Manus. Some have integrated computer‑use agents; others rely on external scaffolds. Two main environments appear: a CLI‑style agent scaffold (OpenHands) and a full computer‑use agent (CUA) using an Ubuntu VM and tools for mouse/keyboard, file editing, and bash. For GPT‑5 they test both scaffolds and, interestingly, the CLI version actually outperforms the CUA in both automation rate and Elo, suggesting current computer‑use agents are not yet extracting value from richer environments. The prompts are deliberately minimalistic but standardized. For non‑computer‑use agents, they instruct the model: read the brief and any inputs, produce exactly the requested deliverables, and package them into a zip. For CLI scaffolds they add explicit instructions about input/output directories and specialized tools: image generation via gpt‑image‑1, TTS via openai/tts‑1, and video via veo‑3.0‑generate‑preview. For the CUA, they define a workspace directory and ask the model to save all deliverables into a Deliverables folder, with helper functions to fix permissions and so on. The evaluation infrastructure is nontrivial. The web‑based platform, as shown in the screenshot on page 23, can natively render a wide range of file types: PDFs, Office docs, Jupyter notebooks, images, video, audio, a variety of CAD formats via Autodesk Viewer, and WebGL/HTML apps. Evaluators are trained with detailed instructions and example failure modes. They’re given soft time caps: 20 minutes per Automation evaluation and 30 minutes per Elo comparison. In practice, Figure 11 shows median times of about 11.6 and 16.9 minutes respectively. Inter‑annotator agreement for the accept/reject decision is 94.4%, which is high given the complexity. So what happens when you unleash these agents? Table 1 and Table 3 give the raw automation rates: Manus tops out at 2.5%, Grok 4 and Sonnet 4.5 at 2.1%, GPT‑5 (CLI) at 1.7%, ChatGPT agent 1.3%, and Gemini 2.5 Pro 0.8%. In other words, out of 240 projects, the best agent cleanly “wins” on roughly six. Table 4 shows the economic side: Manus earns $1,720 out of a possible $143,991; Sonnet 4.5 earns $1,280; GPT‑5 (CLI) $1,180. The rest are lower. So today’s automation impact, on this slice of the market, is economically tiny. Elo scores, plotted in Figure 8, provide more resolution. All agents are far below the human baseline of 1000; the best, Manus, sits around 510, and the others range from ~410 to ~470. Yet differences between models are clearly measurable, and newer models tend to outperform older ones. So while absolute automation is near the floor, relative progress is visible: RLI is sensitive to small quality gains even when tasks remain unsolved. Why are agents failing? The qualitative analysis, summarized in Table 2, clusters failures into four categories: corrupted or unreadable files (17.6% of deliverables), incompleteness (35.7%), poor quality (45.6%), and inconsistencies across assets (14.8%). Examples the paper cites include eight‑second videos where minutes were requested, simplistic geometric “child‑like” drawings, inconsistent 3D renders where a house changes appearance between views, misaligned floorplans versus sketches, or technically functional games with non‑professional graphics. Conversely, the handful of successes occur where current models are strongest: audio production tasks like separating vocals or adding intro/outro music, image‑centric marketing materials—Figure 28 shows a nice Halloween ad example—and code‑only tasks like the interactive World Happiness dashboard in Figure 9. That reinforces the central theme: current LLMs plus tools are powerful in narrow bands, but the remote labor economy is dominated by multimodal, multi‑file, end‑to‑end artifacts where robustness, verification, and aesthetic quality matter. In terms of implications, the authors stress that these results undercut both ungrounded hype and naive extrapolations from code‑or text‑only benchmarks. An automation rate under 3% on a reasonably sampled bundle of freelance work suggests substantial headroom before broad displacement of remote workers. At the same time, the presence of even a small autoflation effect, and clear upward Elo trends, support the idea that as models improve, the effective price of this work will fall—potentially rapidly once certain capability thresholds are crossed. They also argue that unlike domain‑narrow tools such as calculators, AI systems have general cognitive skills and, as Hendrycks et al. put it, are beginning to approximate human‑level cognitive generality. That means a model capable of truly automating RLI likely generalizes to new remote work tasks as they emerge, not just the fixed bundle. The paper is explicit about limitations: many important categories of remote work are excluded by construction—client‑interactive jobs like tutoring, team‑based work like project management, or tasks that can’t be evaluated instantaneously such as SEO. So 100% on RLI would still not imply 100% automation of all remote work. To close, the authors position RLI as an economic analogue of something like MMLU for knowledge: a shared, market‑grounded benchmark to track the trajectory of AI automation. Today, the data say: frontier systems are impressive, but they are not yet remotely close to replacing the broad swath of remote freelancers. How fast that changes is exactly what RLI is designed to tell us.