• Description:

IRC Disentanglement dataset contains over 77,563 messages from Ubuntu IRC channel.

Features include message id, message text and timestamp. Target is list of messages that current message replies to. Each record contains a list of messages from one day of IRC chat.

Split Examples
'test' 10
'train' 153
'validation' 10
  • Feature structure:
    'day': Sequence({
        'id': Text(shape=(), dtype=string),
        'parents': Sequence(Text(shape=(), dtype=string)),
        'text': Text(shape=(), dtype=string),
        'timestamp': Text(shape=(), dtype=string),
  • Feature documentation:
Feature Class Shape Dtype Description
day Sequence
day/id Text string
day/parents Sequence(Text) (None,) string
day/text Text string
day/timestamp Text string
  • Citation:
  author    = {Jonathan K. Kummerfeld and Sai R. Gouravajhala and Joseph Peper and Vignesh Athreya and Chulaka Gunasekara and Jatin Ganhotra and Siva Sankalp Patel and Lazaros Polymenakos and Walter S. Lasecki},
  title     = {A Large-Scale Corpus for Conversation Disentanglement},
  booktitle = {Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics},
  location  = {Florence, Italy},
  month     = {July},
  year      = {2019},
  doi       = {10.18653/v1/P19-1374},
  pages     = {3846--3856},
  url       = {},
  arxiv     = {},
  software  = {},
  data      = {},