Island Parsing

Authors: Luca Ponzanelli, Andrea Mocci

Venue: Shonan Meeting No. 084

Date: 7 March 2016

Abstract

Artifacts containing natural language, like Q&A websites (e.g., Stack Overflow), tutorials, and development emails, are essential to support software development, and they have become a popular subject for software engineering research. The analysis of such artifacts is particularly challenging because of their heterogeneity: these resources consist of natural language interleaved with fragments of multiple programming and markup languages. In this tutorial, I will focus on our efforts towards a systematic approach to model contents of such artifacts, enabling holistic analyses that fully exploit their intrinsic heterogeneous nature. In particular, I will illustrate our StORMeD framework (http://stormed.inf.usi.ch), and how its parsing service can be effectively used to implement a holistic summarizer for Stack Overflow discussions.

Group photo of the participants of NII Shonan Meeting No. 084
NII Shonan Meeting No. 084 — Mining & Modeling Unstructured Data in Software

Island Parsing

7 March 2016 Invited Talk

© Andrea Mocci | CC BY-NC-SA 4.0