AI and software development: a complex debut | ChatGPT 4 | OpenAI free | OpenAI API | Turtles AI

AI and software development: a complex debut
Critical reviews of first autonomous AI engineer show significant performance limitations
Editorial Team23 January 2025

 


An innovative service, presented as the “first AI software engineer,” shows significant limitations in its ability to complete assigned tasks. While demonstrating potential, evaluations show often inefficient and problematic operation.

Key points:

  • Devin was cast as an “AI software engineer” capable of managing projects independently.
  • Evaluation of 20 tasks showed satisfactory completion ability in only 15 percent of cases.
  • Recurring errors include inadequate solutions, technical blockages, and finding nonexistent functions.
  • Testers report a refined but unreliable user experience in actual performance.


In March 2024, the technology landscape was shaken by the announcement of “Devin,” an AI agent dubbed the “first autonomous software engineer.” Created by the Cognition AI organization, the bot promised to improve the software development industry. Capable of writing, testing and executing code, Devin was touted as a comprehensive tool for engineers and development teams, capable of supporting complex projects and even performing personal tasks such as ordering lunch via online platforms. Integrated primarily on Slack, Devin uses a Docker container-based computing environment, with access to tools such as terminals, code editors and external APIs. However, a few months after its commercial launch, set for December 2024 with a starting cost of $500 per month, multiple critical issues are emerging.

Promoted as a solution to automate end-to-end programming, Devin makes use of multiple AI models, including OpenAI’s GPT-4, adapting to new technologies over time. The promise of Cognition AI included operations ranging from code migration to daily assistance tasks. Despite initial enthusiasm, field tests revealed significant inefficiencies. A group of Answer.AI data scientists, consisting of Hamel Husain, Isaac Flath, and Johno Whitaker, put Devin through 20 practice tests and found that only three of them were successfully completed. Successfully completed tasks included extracting data from Notion to import into Google Sheets, creating a planet locator, and researching how to develop a Discord bot in Python. However, for the rest of the challenges, the results were mostly unsatisfactory or inconclusive.

Analysts report how Devin demonstrated increasing difficulty in handling seemingly simple tasks, prolonging them for days and generating unusable or overly complex solutions. A prime example is the attempt to deploy applications on Railway, a platform that does not support such an operation. Devin insisted on unfeasible approaches for over a day, demonstrating little ability to recognize his own limitations. This was compounded by security errors detected by independent developers and even criticism of the validity of some promotional videos released by Cognition AI. Despite an intuitive interface and a well-structured user experience, testers point out the unpredictability of performance and time lost in fruitless attempts. This highlights how the bot’s autonomy, one of its flagship features, has often turned into a hindrance.

Cognition AI has not provided any official statements regarding the issues encountered. Devin, while conceptually promising, leaves open many questions about the future of AI-based software engineers.