Cicirello, Vincent (2020) Kendall tau sequence distance: Extending Kendall tau from ranks to sequences. EAI Endorsed Transactions on Industrial Networks and Intelligent Systems, 7 (23): 1. ISSN 24100218

Text (PDF)
eai.1372018.163925.pdf  Published Version Available under License Creative Commons Attribution No Derivatives. Download (2MB)  Preview 
Abstract
An edit distance is a measure of the minimum cost sequence of edit operations to transform one structure into another. Edit distance can be used as a measure of similarity as part of a pattern recognition system, with lower values of edit distance implying more similar structures. Edit distance is most commonly encountered within the context of strings, where Wagner and Fischer’s string edit distance is perhaps the most wellknown. However, edit distance is not limited to strings. For example, there are several edit distance measures for permutations, including Wagner and Fischer’s string edit distance since a permutation is a special case of a string. However, another edit distance for permutations is Kendall tau distance, which is the number of pairwise element inversions. On permutations, Kendall tau distance is equivalent to an edit distance with adjacent swap as the edit operation. A permutation is often used to represent a total ranking over a set of elements. There exist multiple extensions of Kendall tau distance from total rankings (permutations) to partial rankings (i.e., where multiple elements may have the same rank), but none of these are suitable for computing distance between sequences. We set out to explore extending Kendall tau distance in a different direction, namely from the special case of permutations to the more general case of strings or sequences of elements from some finite alphabet. We name our distance metric Kendall tau sequence distance, and define it as the minimum number of adjacent swaps necessary to transform one sequence into the other. We provide two O(n lg n) algorithms for computing it, and experimentally compare their relative performance. We also provide reference implementations of both algorithms in an open source Java library.
Item Type:  Article 

Uncontrolled Keywords:  edit distance, Kendall tau, pattern recognition, sequences, similarity, strings 
Subjects:  Q Science > QA Mathematics > QA75 Electronic computers. Computer science QA75 Electronic computers. Computer science 
Depositing User:  EAI Editor I. 
Date Deposited:  11 Sep 2020 07:36 
Last Modified:  11 Sep 2020 07:36 
URI:  https://eprints.eudl.eu/id/eprint/220 
Actions (login required)
View Item 