AIS Logo
← Back to Library
RETHINKING KNOWLEDGE WORK: DESIGNING LLM-BASED SYSTEMS FOR COMPLEXITY MANAGEMENT
ECIS 2026 (2026) AI Processed Human Approved

RETHINKING KNOWLEDGE WORK:DESIGNING LLM-BASED SYSTEMS FOR COMPLEXITY MANAGEMENT

Moritz Diener, Simon Kaps, Philipp Spitzer, Robin Hirt, Michael Vössing, Gerhard Satzger
This study applies design science research to design and evaluate an LLM-based system intended to help knowledge workers manage complex, document-intensive workflows. The system is instantiated in a tender analysis tool and evaluated using 34 real-world public procurement documents alongside a qualitative think-aloud study with five domain experts. Problem Knowledge workers face severe cognitive overload due to large, heterogeneous document volumes, strict compliance rules where minor errors lead to immediate disqualification, and complex cross-departmental coordination demands. Prior research fails to offer prescriptive guidelines on designing LLM-based systems to address these multi-dimensional challenges. Outcome - The LLM-based system reduced initial document screening time by 59% (from an average of 3 minutes 42 seconds manually to 1 minute 30 seconds via the system), which was statistically supported.
- The system achieved an overall content-match accuracy of 88.82% for text extraction, with highest accuracy for identifying the procuring entity and place of execution (94.10% each) and lowest for the type of service (79.40%).
- For binary classification tasks, the system achieved an average accuracy of 83.85%, performing significantly better on detecting required certifications (91.20%) than on required references (76.50%).
- Qualitative feedback supported the system's ability to lower perceived document, compliance, and process complexity, shifting human workflows from manual reading and extraction to output verification.
- Limitations include testing within a single German SME in the data center sector, a short-term study duration, and a small evaluation dataset (34 documents), restricting direct generalizability to other industries.
What it means for you
  • CIO / IT Executive: On Monday morning, issue an IT directive mandating that any generative AI tools deployed for document screening must include a 'human-in-the-loop' verification interface, explicitly prohibiting automated final submissions due to the system's average 11.18% error rate in text extraction.
  • IT Manager: On Monday morning, configure your document-processing pipeline's UI to visually highlight and force manual validation of 'type of service' and 'required references' fields, as the research shows these have lower extraction accuracies (79.4% and 76.5%) compared to highly reliable fields like 'procuring entity' (94.1%).
  • Business Strategist: On Monday morning, calculate the potential ROI of deploying an LLM-based screening tool for your proposal-writing team, using the study's proven 59% reduction in initial document screening time (down to 1 minute 30 seconds per document) to project resource hours saved.
  • Researcher: On Monday morning, write a research proposal to extend this design science framework by testing it on a dataset larger than 34 documents across multi-industry and non-German settings, focusing specifically on optimizing prompts to improve the low 76.5% accuracy baseline for identifying required references.
  • Policymaker: On Monday morning, initiate a draft proposal for public procurement formatting standards that mandates standardized, machine-readable sections for 'required certifications' and 'references' to minimize AI extraction errors (currently at 8.8% and 23.5% respectively) and ensure fair bidding processes for SMEs.
Transcript
Host: Welcome to A.I.S. Insights - Turning IS Research into Business Action. I am your host, Aaron Ivan Sterling. Today, we are discussing a study from the European Conference on Information Systems entitled "RETHINKING KNOWLEDGE WORK: DESIGNING LLM-BASED SYSTEMS FOR COMPLEXITY MANAGEMENT". Joining me is our expert analyst, Ava Irene Solis. Welcome, Ava.

Expert: Thank you, Aaron. This study uses a design science research approach to design and evaluate a Large Language Model system, or LLM system, intended to help knowledge workers manage complex, document-intensive workflows, specifically in the context of public tender analysis.

Host: Let's start with the big problem this study addresses. Why is managing document-intensive workflows such a challenge for businesses today?

Expert: The study outlines three main dimensions of complexity that knowledge workers face. First, there is document complexity, which is information overload from navigating massive, non-standardized documents under tight deadlines. Second, we have compliance complexity, where missing a single required certificate or detail can lead to immediate disqualification. Finally, there is process complexity, which involves coordination demands across sales, legal, and technical departments. Currently, prior research lacks prescriptive guidelines on how to design LLM-based systems to address these three overlapping challenges.

Host: How did the researchers approach this problem to find a solution?

Expert: They collaborated with a German small-to-medium enterprise, or SME, in the data center sector that regularly bids on public tenders. By interviewing domain experts, they derived design requirements and instantiated them into an LLM-based tender analysis tool. This tool features a structured tender overview, a compliance checklist and timeline, and a source-grounded question-and-answering feature. They evaluated this tool quantitatively using thirty-four real-world public procurement documents and qualitatively through a think-aloud study with five domain experts.

Host: What were the key findings from this evaluation?

Expert: Statistically, the system showed a significant reduction in initial screening time. On average, manual screening by experts took three minutes and forty-two seconds per document. Using the LLM-based system, that dropped to one minute and thirty seconds, which is a fifty-nine percent time saving.

Host: That is a substantial decrease in time. How did the system perform in terms of accuracy?

Expert: For text extraction, the system achieved an overall content-match accuracy of eighty-eight point eight-two percent. However, the accuracy varied depending on the item. It was highest for identifying the procuring entity and place of execution at ninety-four point ten percent each, but lower for extracting the type of service, which was seventy-nine point four-zero percent. This is because service descriptions are often written in highly heterogeneous, domain-specific language.

Host: And what about the binary tasks, like checking if certain documents were required?

Expert: For yes-or-no classification tasks, the average accuracy was eighty-three point eight-five percent. It performed better on detecting required certifications at ninety-one point two-zero percent than on identifying required references, which was seventy-six point five-zero percent. This difference occurred because references are often described indirectly or as optional elements, whereas certifications are usually explicitly stated.

Host: Did the study point out any limitations to these findings?

Expert: Yes, the evaluation was conducted with a single German SME in the data center sector over a short-term period using a small dataset of thirty-four documents. Therefore, the findings may not directly generalize to other industries or larger organizational settings without further evaluation.

Host: Knowing this, how can business and technology practitioners use this study in their operations? What is the major takeaway?

Expert: The key takeaway is that we should reposition GenAI not as a tool for complete, autonomous automation, but rather as a system for complexity management and redistribution. The qualitative feedback showed that the system shifted the human workflow from manual reading and extraction to what the researchers call an "orchestrate-and-verify" mode. Instead of reading hundreds of pages, humans orchestrate the AI to retrieve summaries and then use linked sources to verify the output.

Host: So, it redefines the skills our workforce actually needs?

Expert: Exactly. Instead of training employees solely to read and manually extract data, organizations need to build AI literacy. This means training workers on how to construct effective requests, interpret confidence signals, understand system limitations, and systematically verify the AI's output using source links. The study suggests that organizations that invest in these hybrid competencies will be better positioned to leverage LLMs while avoiding compliance risks and potential long-term deskilling.

Host: That is a highly practical perspective on integrating AI into complex business workflows. Thank you, Ava, for sharing these insights with us today.

Expert: It was a pleasure, Aaron.

Host: And thank you to our listeners for tuning in to A.I.S. Insights. Join us next time as we continue to translate information systems research into actionable business strategies.
large language models, generative AI, complexity management, knowledge work, design science research