results highlight the importance of previously overlooked design choices, and raise questions about the sourceThe original BERT uses a subword-level tokenization with the vocabulary size of 30K which is learned after input preprocessing and using several heuristics. RoBERTa uses bytes instea