AIS Logo
← Back to Library
What Does ChatGPT Know About Information Systems?
Communications of the Association for Information Systems (2025) AI Processed Human Approved

What Does ChatGPT Know About Information Systems?

Daniel E. O’Leary, Veda C. Storey, Aaron M. French, Joseph R. Buckman, Cecil Chua, Andrew W. Green, Grace Gu, Fred Niederman, Francis Pereira, Gary Templeton, Linda Wallace
This study evaluates the depth and quality of domain-specific knowledge in ChatGPT (GPT-3.5) within the field of Information Systems (IS). Using an established IS competency framework, the researchers analyzed ChatGPT's performance on over 3,000 questions gathered from university IS courses and professional certification exams. Problem While Generative AI tools demonstrate impressive general academic abilities, their specific knowledge accuracy across specialized fields like Information Systems remains uncertain. Educators and industry professionals lack clear benchmarks regarding whether ChatGPT can reliably produce accurate IS content, pass professional certifications, or answer complex technical queries without hallucinating. Outcome - ChatGPT answered 65% to 85% of Information Systems queries correctly across all core framework areas, performing roughly at the level of an average student.
- The model performed significantly better on university course exams than on professional certification exams, failing to reach passing scores for the PMP and Microsoft Azure Administrator tests.
- ChatGPT achieved higher relative accuracy against class averages on closed-book exams compared to open-book or online exam environments where students had external resource access.
- While ChatGPT displayed broad coverage across diverse IS topics, its responses favored listing structured information over demonstrating expert-level critical thinking or complex problem-solving.
What it means for you
  • CIO / IT Executive: On Monday morning, task your IT training lead to pilot a small internal workshop for a select group of IT professionals focusing on the PMP and Microsoft Azure Administrator certifications, using ChatGPT as a supplementary learning tool to identify specific knowledge gaps and assess its effectiveness in preparing for these professional exams.
  • IT Manager: On Monday morning, instruct your team members responsible for generating internal documentation or knowledge base articles to compare the accuracy and completeness of ChatGPT's initial drafts on at least two core Information Systems topics (e.g., database management, network security) against your existing, verified internal resources.
  • Business Strategist: On Monday morning, task a market analyst to explore potential use cases for ChatGPT in rapidly generating initial drafts of business process documentation or competitive analysis reports within your IS domain, noting its tendency to provide structured information over deep critical thinking.
  • Researcher: On Monday morning, design a controlled experiment to compare ChatGPT's performance on retrieving specific, factual information about IS concepts in a closed-book scenario versus an open-book scenario using publicly available academic papers, to validate the finding about its performance in different exam environments.
  • Policymaker: On Monday morning, request a briefing from your education and technology policy advisors on the implications of AI tools like ChatGPT achieving 'average student' performance in IS education, and explore potential guidelines for incorporating such tools responsibly in curriculum development and assessment.
Transcript
Host: Welcome to A.I.S. Insights — powered by Living Knowledge. I'm Anna Ivy Summers. Today, we're diving into a question that every business leader, tech manager, and educator is asking right now: what does Generative AI actually know when it comes to specialized fields? Specifically, we're breaking down a major study titled "What Does ChatGPT Know About Information Systems?" Joining me to explain the details is our lead analyst, Alex Ian Sutherland. Welcome, Alex!

Expert: Thanks, Anna. It's great to be here.

Host: So, Alex, we see ChatGPT writing emails, drafting code, and summarizing articles every day. But when it comes to a specialized, fast-moving field like Information Systems—or IS—how deep is its actual domain knowledge? What prompted researchers to investigate this?

Expert: That's the core problem, Anna. Large language models like ChatGPT are trained on vast amounts of general internet text, creating an "average" map of language across many subjects. But Information Systems is a rapidly evolving discipline that requires real technical depth. Because these models are based on complex neural networks, they operate as a "black box." We can't simply look inside to see what facts they know or don't know. If the model encounters something outside its training, it doesn't usually admit it—it hallucinates an answer. So, without rigorous testing, businesses and educators can't know whether they can actually depend on ChatGPT for technical accuracy.

Host: That makes total sense. We need a clear benchmark. How did the research team design this study to test ChatGPT's knowledge?

Expert: The researchers used a well-established framework for Information Systems education that covers six core areas, including data management, IT infrastructure, development, project management, and organizational strategy. They gathered over 3,000 exam and quiz questions from ten faculty members across seven different universities. On top of that, they included questions from industry professional certification exams, like the PMP for project management, CISSP for cybersecurity, and Microsoft Azure Administrator. They submitted all these questions to ChatGPT and analyzed its accuracy down to the item level.

Host: Over 3,000 questions across both academic and professional levels—that's a comprehensive approach. What were the key findings?

Expert: Across all the core areas of Information Systems, ChatGPT scored between 65% and 85% correct. Overall, its performance matched roughly the level of an average undergraduate or graduate student. What was particularly interesting is that its knowledge was very consistent across the different IS domains—whether it was database concepts or project integration, it performed at a similar baseline.

Host: That sounds reasonably strong for general questions, but how did it perform on those high-stakes professional certification exams?

Expert: That's where we see a clear limitation. ChatGPT performed significantly worse on professional exams than on university course tests. For instance, it scored under 70% on practice exams for the Project Management Professional certification and 74% on the Microsoft Azure Administrator test, failing both. It did better on the CISSP security exam with about an 80% score, but overall, the study showed that ChatGPT is not operating at an expert practitioner level.

Host: That is a crucial distinction for businesses. Did the study look at different question formats, like multiple-choice versus essay questions?

Expert: Yes, and the results were surprising at first glance. ChatGPT actually scored higher on essay questions, around 88%, compared to multiple-choice at 81%. However, when the researchers analyzed why, they found that most essay questions asked the model to "list" or explain factors. ChatGPT excels at generating structured lists, but it struggled to demonstrate true critical thinking or complex logical arguments. Another telling result was that ChatGPT outperformed class averages much more frequently on closed-book exams than on open-book or online exams, where human students had access to external resources.

Host: That brings us to the bottom line for business professionals. What are the key takeaways from this study for organizations and leaders looking to leverage AI?

Expert: The biggest takeaway is that while ChatGPT is a remarkably broad generalist, it should not be treated as an unquestioned domain expert. For business leaders, relying on an out-of-the-box LLM for complex IT architecture, project execution, or compliance decisions carries real risk. That's why many organizations are supplementing LLMs with retrieval-augmented generation—or RAG—and structured prompt management to ground the AI in verified corporate knowledge. For educators and trainers, the study shows that assessment needs to shift away from simple list-building questions toward critical thinking and real-world case studies, where human expertise still clearly outperforms AI.

Host: A great reminder that while AI is an impressive tool, human oversight and critical thinking remain essential. Alex, thank you so much for breaking down this study for us today.

Expert: My pleasure, Anna.

Host: And thank you for listening to A.I.S. Insights — powered by Living Knowledge. Be sure to subscribe and join us next time for more insights at the intersection of technology and business strategy. Until then, stay curious!
ChatGPT, Large Language Models, Generative AI, Information Systems Competency, Artificial Intelligence in Education, Knowledge Assessment