Abstract
Objective: The minimum transformer performance can move forward by incorporating the progress of the last decade. Background: Current efforts begin with surface forms and build meanings for tagging, parsing, identifying semantic relationships, etc. But a parser and outside resources can also provide these morphological, syntactic, and semantic features. Method: Instead of inputting tokens of surface forms into a transformer, we could start with feature vectors whose summed embeddings represent the deep features of the words, changing it to structured data. This is an empirical study on how inputs with feature vectors perform on masked word prediction. Compared to classic transformers, increasing integration between the parser and transformer accumulates to a total accuracy gain of 143% and perplexity loss of 87%.


