StORMeD: Stack Overflow Ready Made Data

StORMeD: Stack Overflow Ready Made Data

16 May 2015 Paper

Authors: Luca Ponzanelli, Andrea Mocci, Michele Lanza

Proceedings of MSR 2015 (12th Working Conference on Mining Software Repositories, Data Showcase Track)

Citations
Google Scholar

Abstract. Stack Overflow is the de facto Question and Answer (Q&A) website for developers, and it has been used in many approaches by software engineering researchers to mine useful data. However, the contents of a Stack Overflow discussion are inherently heterogeneous, mixing natural language, source code, stack traces and configuration files in XML or JSON format. We constructed a full island grammar capable of modeling the set of 700,000 Stack Overflow discussions talking about Java, building a heterogeneous abstract syntax tree (H-AST) of each post (question, answer or comment) in a discussion. The resulting dataset models every Stack Overflow discussion, providing a full H-AST for each type of structured fragment (i.e., JSON, XML, Java, Stack traces), and complementing this information with a set of basic meta-information like term frequency to enable natural language analyses. Our dataset allows the end-user to perform combined analyses of the Stack Overflow by visiting the H-AST of a discussion.

Coauthors

Pre-Print

Project

1 April 2014 Project Postdoctoral Researcher

ESSENTIALS: People-centric Essentials for Software Evolution

Shifting the focus of software evolution research to the people-centric 'evolutionary essentials' that stakeholders need in their current working context.

ESSENTIALS: People-centric Essentials for Software Evolution

Thanks for reading!

StORMeD: Stack Overflow Ready Made Data

16 May 2015 Paper 0