dorsal/arxiv
View SchemaEvaluating Novelty in AI-Generated Research Plans Using Multi-Workflow LLM Pipelines
| Authors | Devesh Saraogi, Rohit Singhee, Dhruv Kumar |
|---|---|
| Categories | |
| ArXiv ID | 2601.09714vv1 |
| URL | https://arxiv.org/abs/2601.09714 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
The integration of Large Language Models (LLMs) into the scientific ecosystem raises fundamental questions about the creativity and originality of AI-generated research. Recent work has identified ``smart plagiarism'' as a concern in single-step prompting approaches, where models reproduce existing ideas with terminological shifts. This paper investigates whether agentic workflows -- multi-step systems employing iterative reasoning, evolutionary search, and recursive decomposition -- can generate more novel and feasible research plans. We benchmark five reasoning architectures: Reflection-based iterative refinement, Sakana AI v2 evolutionary algorithms, Google Co-Scientist multi-agent framework, GPT Deep Research (GPT-5.1) recursive decomposition, and Gemini~3 Pro multimodal long-context pipeline. Using evaluations from thirty proposals each on novelty, feasibility, and impact, we find that decomposition-based and long-context workflows achieve mean novelty of 4.17/5, while reflection-based approaches score significantly lower (2.33/5). Results reveal varied performance across research domains, with high-performing workflows maintaining feasibility without sacrificing creativity. These findings support the view that carefully designed multi-stage agentic workflows can advance AI-assisted research ideation.
{
"annotation_id": "ffa3baa4-461b-48fb-b2ea-112de1a10a13",
"date_created": "2026-02-17T05:53:20.453000Z",
"date_modified": "2026-02-17T05:53:20.453000Z",
"file_hash": "d7addd34ff3d8b93809652e55b04e1da159075858a85220de54e5fab13d2a44b",
"private": false,
"record": {
"abstract": "The integration of Large Language Models (LLMs) into the scientific ecosystem raises fundamental questions about the creativity and originality of AI-generated research. Recent work has identified ``smart plagiarism\u0027\u0027 as a concern in single-step prompting approaches, where models reproduce existing ideas with terminological shifts. This paper investigates whether agentic workflows -- multi-step systems employing iterative reasoning, evolutionary search, and recursive decomposition -- can generate more novel and feasible research plans. We benchmark five reasoning architectures: Reflection-based iterative refinement, Sakana AI v2 evolutionary algorithms, Google Co-Scientist multi-agent framework, GPT Deep Research (GPT-5.1) recursive decomposition, and Gemini~3 Pro multimodal long-context pipeline. Using evaluations from thirty proposals each on novelty, feasibility, and impact, we find that decomposition-based and long-context workflows achieve mean novelty of 4.17/5, while reflection-based approaches score significantly lower (2.33/5). Results reveal varied performance across research domains, with high-performing workflows maintaining feasibility without sacrificing creativity. These findings support the view that carefully designed multi-stage agentic workflows can advance AI-assisted research ideation.",
"arxiv_id": "2601.09714",
"authors": [
"Devesh Saraogi",
"Rohit Singhee",
"Dhruv Kumar"
],
"categories": [
"cs.CL",
"cs.AI"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "Evaluating Novelty in AI-Generated Research Plans Using Multi-Workflow LLM Pipelines",
"url": "https://arxiv.org/abs/2601.09714",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "c412cf1e-8ac8-49f1-8e5d-04e205bbf323",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}