From HandWiki - Reading time: 2 min
This article relies largely or entirely on a single source. (May 2026) |
Software mining is a subfield of software engineering that focuses on extracting and analyzing information from software artifacts stored in repositories such as version control systems, issue trackers, and communication logs. It aims to uncover patterns and actionable insights about software systems and development processes using techniques such as data mining, statistical analysis, and machine learning, supporting activities like software maintenance, evolution, and quality assessment.[1]
Developed specification Knowledge Discovery Metamodel (KDM) which defines an ontology for software assets and their relationships for the purpose of performing knowledge discovery of existing code. The OMG Knowledge Discovery Metamodel provides an integrated representation to capturing application metadata. Another OMG specification, the Common Warehouse Metamodel focuses entirely on mining enterprise metadata.
Software mining is closely related to data mining, since existing software artifacts contain enormous business value, key for the evolution of software systems. Knowledge discovery from software systems addresses structure, behavior as well as the data processed by the software system. Instead of mining individual data sets, software mining focuses on metadata.
Text mining software tools enable easy handling of text documents for the purpose of data analysis including automatic model generation and document classification, document clustering, document visualization, dealing with Web documents, and crawling the Web.
Knowledge discovery in software is related to a concept of reverse engineering. Software mining addresses structure, behavior as well as the data processed by the software system.
Mining software systems may happen at various levels: