Results 1 -
2 of
2
Evaluating Natural Language Processing Systems
, 1993
"... This report presents a detailed analysis and review of NLP evaluation, in principle and in practice. Part 1 examines evaluation concepts and establishes a framework for NLP system evaluation. This makes use of experience in the related area of information retrieval and the analysis also refers to ev ..."
Abstract
-
Cited by 104 (0 self)
- Add to MetaCart
This report presents a detailed analysis and review of NLP evaluation, in principle and in practice. Part 1 examines evaluation concepts and establishes a framework for NLP system evaluation. This makes use of experience in the related area of information retrieval and the analysis also refers to evaluation in speech processing. Part 2 surveys significant evaluation work done so far, for instance in machine translation, and discusses the particular problems of generic system evaluation. The conclusion is that evaluation strategies and techniques for NLP need much more development, in particular to take proper account of the influence of system tasks and settings. Part 3 develops a general approach to NLP evaluation, aimed at methodologically-sound strategies for test and evaluation motivated by comprehensive performance factor identification. The analysis throughout the report is supported by extensive illustrative examples. This work was carried out under the UK Science and Engineeri...
Contents
"... this report . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7 1.2.2 How to read it . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7 1.3 Types of systems considered . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 9 1.3.1 Long term . . . . . . ..."
Abstract
- Add to MetaCart
this report . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7 1.2.2 How to read it . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7 1.3 Types of systems considered . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 9 1.3.1 Long term . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 9 1.3.2 Short term . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 10 2 The Framework Model 11 2.1 The ISO 9126 Standard . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11 2.1.1 ISO 9126 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11 2.1.1.1 Quality requirements definition . . . . . . . . . . . . . . . . . . . . . 12 2.1.1.2 Evaluation preparation . . . . . . . . . . . . . . . . . . . . . . . . . 12 2.1.1.3 Evaluation procedure . . . . . . . . . . . . . . . . . . . . . . . . . . 13 2.1.2 The EAGLES extensions to ISO 9126 . . . . . . . . . . . . . . . . . . . . . . 13 2.2 Towards formalisation and automation . . . . . . . . . . . . . . . . . . . . . . . . . . 15 2.2.1 Key concepts in evaluation --- a sketch for a formalisation . . . . . . . . . . . 15 2.2.1.1 The evaluation function . . . . . . . . . . . . . . . . . . . . . . . . . 15 2.2.1.2 Feature descriptions . . . . . . . . . . . . . . . . . . . . . . . . . . . 16 2.2.1.3 Some useful terms . . . . . . . . . . . . . . . . . . . . . . . . . . . . 17 2.2.1.4 Candidates for standardisation . . . . . . . . . . . . . . . . . . . . . 18 2.2.2 Parameterisable test bed . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 18 2.2.2.1 Parameterisable Test Bed . . . . . . . . . . . . . . . . . . . . . . . . 19 2.2.2.1.1 Parameters of objects . . . . . . . . . . . . . . . . . . . . . 19 2...

