Authors: Luca Ponzanelli, Andrea Mocci
Venue: Shonan Meeting No. 084
Date: 7 March 2016

Authors: Luca Ponzanelli, Andrea Mocci
Venue: Shonan Meeting No. 084
Date: 7 March 2016
Artifacts containing natural language, like Q&A websites (e.g., Stack Overflow), tutorials, and development emails, are essential to support software development, and they have become a popular subject for software engineering research. The analysis of such artifacts is particularly challenging because of their heterogeneity: these resources consist of natural language interleaved with fragments of multiple programming and markup languages. In this tutorial, I will focus on our efforts towards a systematic approach to model contents of such artifacts, enabling holistic analyses that fully exploit their intrinsic heterogeneous nature. In particular, I will illustrate our StORMeD framework (http://stormed.inf.usi.ch↗), and how its parsing service can be effectively used to implement a holistic summarizer for Stack Overflow discussions.
