keyword
generative AI
Generative artificial intelligence refers to a class of artificial intelligence technologies capable of creating new content—such as text, computer code, images, audio, and video—based on patterns and structures learned from training data. Unlike traditional discriminative systems that primarily classify, recognize, or analyze existing information, generative models leverage deep learning architectures, such as transformer-based large language models and diffusion models, to synthesize novel, contextually relevant outputs in response to user prompts. These systems are widely utilized across domains including software engineering, automated text processing, multimedia creation, and complex reasoning tasks, serving to augment human productivity while requiring ongoing evaluation, alignment, and oversight to ensure reliability and safety.
7 items

To Copilot and Beyond: 22 AI Systems Developers Want Built
Rudrajit Choudhuri, Christian Bird, Carmen Badea, Anita Sarma
Why you should read this
Identifies 22 AI systems software engineers want built beyond code generation, using a survey of 860 developers to establish the principle of bounded delegation for designing tools that offload peripheral tasks without intruding on core engineering craft.
Developers spend roughly one-tenth of their workday writing code, yet most AI tooling targets that fraction. This paper asks what should be built for the rest. We surveyed 860 Microsoft developers to understand where they want AI support, and where they want it to stay out. Using a human-in-the-loop, multi-model council-based thematic analysis, we identify 22 AI systems that developers want built across five task categories. For each, we describe the problem it solves, what makes it hard to build, and the constraints developers place on its behavior. Our findings point to a growing right-shift burden in AI-assisted development: developers wanted systems that embed quality signals earlier in their workflow to keep pace with accelerating code generation, while enforcing explicit authority scoping, provenance, uncertainty signaling, and least-privilege access throughout. This tension reveals a pattern we call "bounded delegation": developers wanted AI to absorb the assembly work surrounding their craft, never the craft itself. That boundary tracks where they locate professional identity, suggesting that the value of AI tooling may lie as much in where and how precisely it stops as in what it does.
Added
2026-09-30

AI Where It Matters: Where, Why, and How Developers Want AI Support in Daily Work
Rudrajit Choudhuri, Carmen Badea, Christian Bird, Jenna Butler, Robert DeLIne, Brian Houck
Why you should read this
Presents an empirical model derived from 860 developers that uses cognitive appraisal theory to identify where practitioners want generative AI assistance, where they resist automation, and how responsible AI priorities shift across software engineering tasks.
Generative AI is reshaping software work, yet we lack clear guidance on where developers most need support and how to design it responsibly. We report a large-scale, mixed-methods study of N=860 developers examining where, why, and how they seek or limit AI help across SE tasks. Using cognitive appraisal theory, we provide the first empirically validated mapping of developers' task appraisals to AI adoption patterns and Responsible AI (RAI) priorities. Appraisals predict AI openness and use, revealing distinct patterns: strong current use and demand for improvement in core work (e.g., coding, testing); high demand to reduce toil (e.g., documentation, operations); and clear limits for identity- and relationship-centric work (e.g., mentoring). RAI priorities vary by context: reliability and security for systems-facing tasks; transparency, alignment, and steerability to maintain control; and fairness and inclusiveness for human-facing work. Our results offer concrete, contextual guidance for delivering AI where it matters to developers and their work.
Added
2026-09-29

Keeping Humans in the Loop: Human-Centered Automated Annotation with Generative AI
Nick Pangakis, Sam Wolken
Why you should read this
Demonstrates through twenty-seven private social science tasks that large language model annotations vary unpredictably and diverge from human judgment, proving that human validation remains essential for automated research workflows.
Automated text annotation is a compelling use case for generative large language models (LLMs) in social media research. Recent work suggests that LLMs can achieve strong performance on annotation tasks; however, these studies evaluate LLMs on a small number of tasks and likely suffer from contamination due to a reliance on public benchmark datasets. Here, we test a human-centered framework for responsibly evaluating artificial intelligence tools used in automated annotation. We use GPT-4 to replicate 27 annotation tasks across 11 password-protected datasets from recently published computational social science articles in high-impact journals. For each task, we compare GPT-4 annotations against human-annotated ground-truth labels and against annotations from separate supervised classification models fine-tuned on human-generated labels. Although the quality of LLM labels is generally high, we find significant variation in LLM performance across tasks, even within datasets. Our findings underscore the importance of a human-centered workflow and careful evaluation standards: Automated annotations significantly diverge from human judgment in numerous scenarios, despite various optimization strategies such as prompt tuning. Grounding automated annotation in validation labels generated by humans is essential for responsible evaluation.
Added
2026-09-29

The SPACE of AI: Real-World Lessons on AI's Impact on Developers
Brian Houck, Travis Lowdermilk, Cody Beyer, Steven Clarke, Ben Hanrahan
Why you should read this
Demonstrates how AI alters developer productivity across the SPACE framework, drawing on data from over 500 software engineers to show that efficiency gains in routine tasks depend directly on team culture and organizational support.
As artificial intelligence (AI) tools become increasingly embedded in software development workflows, questions persist about their true impact on developer productivity and experience. This paper presents findings from a mixed-methods study examining how developers perceive AI's influence across the dimensions of the SPACE framework: Satisfaction, Performance, Activity, Collaboration and Efficiency. Drawing on survey responses from over 500 developers and qualitative insights from interviews and observational studies, we find that AI is broadly adopted and widely seen as enhancing productivity, particularly for routine tasks. However, the benefits vary, depending on task complexity, individual usage patterns, and team-level adoption. Developers report increased efficiency and satisfaction, with less evidence of impact on collaboration. Organizational support and peer learning play key roles in maximizing AI's value. These findings suggest that AI is augmenting developers rather than replacing them, and that effective integration depends as much on team culture and support structures as on the tools themselves. We conclude with practical recommendations for teams, organizations and researchers seeking to harness AI's potential in software engineering.
Added
2026-09-29

MEGA: Multilingual Evaluation of Generative AI
Kabir Ahuja, Harshita Diddee, Rishav Hada, Millicent Ochieng, Krithika Ramesh, Prachi Jain, Akshay Uttama Nambi, Tanuja Ganu, Sameer Segal, Mohamed Ahmed, Kalika Bali, Sunayana Sitaram
Why you should read this
Introduces MEGA, a comprehensive multilingual benchmark spanning 16 standard tasks across 70 typologically diverse languages to systematically compare generative large language models against prior state-of-the-art systems and identify critical performance gaps in low-resource settings.
Generative AI models have shown impressive performance on many Natural Language Processing tasks such as language understanding, reasoning, and language generation. An important question being asked by the AI community today is about the capabilities and limits of these models, and it is clear that evaluating generative AI is very challenging. Most studies on generative LLMs have been restricted to English and it is unclear how capable these models are at understanding and generating text in other languages. We present the first comprehensive benchmarking of generative LLMs - MEGA, which evaluates models on standard NLP benchmarks, covering 16 NLP datasets across 70 typologically diverse languages. We compare the performance of generative LLMs including Chat-GPT and GPT-4 to State of the Art (SOTA) non-autoregressive models on these tasks to determine how well generative models perform compared to the previous generation of LLMs. We present a thorough analysis of the performance of models across languages and tasks and discuss challenges in improving the performance of generative LLMs on low-resource languages. We create a framework for evaluating generative LLMs in the multilingual setting and provide directions for future progress in the field.
Added
2026-09-26

FilMaster: Bridging Cinematic Principles and Generative AI for Automated Film Generation
Kaiyi Huang, Yukun Huang, Xintao Wang, Zinan Lin, Xuefei Ning, Pengfei Wan, Di Zhang, Yu Wang, Xihui Liu
Why you should read this
Introduces AGUVIS, a unified vision-based framework that achieves state-of-the-art autonomous GUI interaction across web, desktop, and mobile platforms by combining direct screen perception with a structured inner monologue, establishing the first high-performance open-source alternative to agents dependent on closed-source models.
AI-driven content creation has shown potential in film production. However, existing film generation systems struggle to implement cinematic principles and thus fail to generate professional-quality films, particularly lacking diverse camera language and cinematic rhythm. This results in templated visuals and unengaging narratives. To address this, we introduce FilMaster, an end-to-end AI system that integrates real-world cinematic principles for professional-grade film generation, yielding editable, industry-standard outputs. FilMaster is built on two key principles: (1) learning cinematography from extensive real-world film data and (2) emulating professional, audience-centric post-production workflows. Inspired by these principles, FilMaster incorporates two stages: a Reference-Guided Generation Stage which transforms user input to video clips, and a Generative Post-Production Stage which transforms raw footage into audiovisual outputs by orchestrating visual and auditory elements for cinematic rhythm. Our generation stage highlights a Multi-shot Synergized RAG Camera Language Design module to guide the AI in generating professional camera language by retrieving reference clips from a vast corpus of 440,000 film clips. Our post-production stage emulates professional workflows by designing an Audience-Centric Cinematic Rhythm Control module, including Rough Cut and Fine Cut processes informed by simulated audience feedback, for effective integration of audiovisual elements to achieve engaging content. The system is empowered by generative AI models like (M)LLMs and video generation models. Furthermore, we introduce FilmEval, a comprehensive benchmark for evaluating AI-generated films. Extensive experiments show FilMaster's superior performance in camera language design and cinematic rhythm control, advancing generative AI in professional filmmaking.
Added
2026-05-16

Brief analysis of DeepSeek R1 and its implications for Generative AI
Sarah Mercer, Samuel Spillard, Daniel P. Martin
Why you should read this
Explains how DeepSeek R1 achieves competitive performance against OpenAI's models at a fraction of the cost, making it essential reading for anyone interested in cost-effective advancements in Generative AI.
In late January 2025, DeepSeek released their new reasoning model (DeepSeek R1); which was developed at a fraction of the cost yet remains competitive with OpenAI's models, despite the US's GPU export ban. This report discusses the model, and what its release means for the field of Generative AI more widely. We briefly discuss other models released from China in recent weeks, their similarities; innovative use of Mixture of Experts (MoE), Reinforcement Learning (RL) and clever engineering appear to be key factors in the capabilities of these models. This think piece has been written to a tight timescale, providing broad coverage of the topic, and serves as introductory material for those looking to understand the model's technical advancements, as well as its place in the ecosystem. Several further areas of research are identified.
Added
2026-01-09

