| ESP Investigación cuantitativa |
contents
Introduction - A snapshot of quantitative research | The blurred boundaries of quantitative vs qualitative research | The quantitative research design | Research potential
Introduction - A snapshot of quantitative research
Quantitative research is usually defined in very simple terms as a type of research dealing with numerical data or at least with data that can be converted into numbers (Sheard 2018). This numerical nature of data is frequently associated with scales of measurement and analysis through statistical modelling. Data-focused information—namely, the nature of data and the procedures for its collection, measurement and analysis—becomes the major parameter guiding definitions of quantitative research.
Besides data-focused attributes, there are three key features related to the scientific method, which are generally ascribed to quantitative research. Two refer to essential conditions of hypothesis testing, namely the use of a hypothetico-deductive method that needs to be falsifiable. According to these conditions, predictions in quantitative research should be formally stated in a way that can be falsifiable, that is, proven to be wrong by objective testing. The third feature refers to the need for replicability, viz., for the possibility to test the reported findings by repeating the research method. Replicability is related to two traits also expected from quantitative researchers’ conduct: honesty in the way the research has been carried out and confidence that the results obtained have been the best possible. Replicability involves the use of the same analysis with different data, and should not be confused with reproducibility (same data and same analysis), nor with the robustness (same data, different analysis) or generalization (different data, different analysis) of the data.
![]() |
In the literature on research methods, quantitative and qualitative research are usually defined in tandem, with one serving to illustrate the opposite of the other, but with a general desideratum for them to work together. In education, health sciences, sociology, psychology and also in translation and interpreting, a third methodological type is also introduced along with quantitative and qualitative research, i.e., mixed methods research, which integrates and combines the other two (e.g., Saldanha & O’Brien 2013)
Within Translation and Interpreting Studies (TIS), the research upsurge in Cognitive Translation & Interpreting Studies (CTIS) in the 1990s towards a more scientific and empirical methodology favored the expansion of quantitative research at the expense of qualitative investigation (Ferreira, Schwieter & Gile 2015: 4). In CTIS, the shift came hand in hand with a surge of experimental designs on translators' cognitive processes, which have been criticized, among other reasons, for the scarce ecological validity of their settings—especially when the focus was on professional translation—and the lack of information on the quality and nature of data. As a result, TIS research has veered into mixed methods approaches as a way to achieve a deeper understanding of the research problem and a more accurate approximation to the examined real-world professional and learning populations.
Quantitative and qualitative methods certainly display some noticeable differences in their approaches to the key parameters for a research design. Nonetheless, the neat and clear boundaries created to separate them are patently artificial; in TIS boundaries are blurred and the distance between them is not as large as it seems (Hale & Napier 2013). Rather, research is seen as occurring along a continuum with designs that may be more or less quantitatively or qualitatively oriented, but very rarely positioned at the far ends (Rojo 2013). Despite the soft boundaries, and for the sake of theoretical description, here they are discussed discreetly, and the present entry provides a snapshot of quantitative research that profiles the lack of neat boundaries between the categories and discusses some of the perks and perils of quantitative research with regard to the basic components of a quantitative research design, i.e., the related studies and existing gaps in the literature; research questions and variables; hypotheses, sample size and inference tests; measurement tools; and the analysis and report of results. We take a quick glimpse into the future of responsible research in TIS.
The blurred boundaries of quantitative vs qualitative research
Definitions of quantitative and qualitative research commonly portray them as a clear-cut, multi-attribute dichotomy in which both approaches are assumed to be polar opposites competing with each other. Over the history of TIS, this rivalry has been reflected in the intermittent shift in the balance of power towards one or the other, very much depending on the theoretical approach (Rojo 2013). And yet, the progressive realization of the limitations of single method approaches has finally led both parties to settle for a ‘technical draw’, to the advantage of mixed methods approaches.
Over the last decades, there has been in fact a growing appreciation in TIS for mixed methods research, partly motivated by the need to overcome the limitations of a single method approach, and partly by the realization that quantitative and qualitative approaches represent in practice an interactive continuum. Most studies—especially in the humanities and social sciences—involve both quantitative and qualitative assumptions to at least some degree (Ridenour & Newman 2008: 17). A clear case of this spongy boundary condition is the fact that a big variety of qualitative information can be converted into quantitative data amenable to be computed and statistically analyzed—the most typical case in TIS being probably that of translation problems and strategies (Rojo 2019). The majority of researchers would, therefore, advocate for more flexible mixed paradigms allowing them to make the most of quantitative and qualitative methods depending on their specific research needs.
In mixed methods designs, the quantitative approach may either precede or follow the qualitative method, depending on which one is given priority and how data collection is implemented. As indicated, quantitative methods are favored when quantity or numbers are emphasized, i.e., when the variables involved in the research question or hypothesis can be quantified, measured and tested with a large sample of the target population and/or data. In combination with qualitative methods, quantitative ones may serve to provide numerical support to a previous in-depth exploration by selecting those data with potential to be applied to a larger population or, alternatively, may initially serve to test variables with a large sample before a few cases are later on explored in greater depth.
Adding qualitative insights to quantitative data is particularly appropriate in disciplines like TIS, where language is of special significance for the study of human interaction in social contexts, and where numbers may on their own be insufficient to account for the richness and complexity of human communication. In what follows, we will describe the key parameters of quantitative methodology, discussing some of their potential pros and cons and suggesting ways to overcome the latter in the future.
The quantitative research design
The labels research and methodology have so far been used rather interchangeably to refer to the general framework or overall strategy guiding a study. This strategy—a blueprint for collecting, measuring and analyzing data—splits into three methods: quantitative, qualitative and mixed.
Let us focus on the design. A research design can be defined as the specific outline describing how the selected method is applied to answer a particular research question, by integrating the components in the study into a coherent and logical whole. The design is particularly relevant in quantitative methodology, which typically follows a closed and predictable structure planned in advance. The description of a prototypical quantitative design has been structured into five different stages: identification of existing gaps in the literature; choice of the research question and variables; formulation of hypotheses and the relevance of sample size and inference tests; selection of adequate measurement tools; and analysis and report of results.
Identifying the gaps in the literature
The first step to build a research design is defining the problem we want to investigate. The choice of a research problem may be guided by many reasons, from personal curiosity to the desire to explore a theory. Whatever the motivation, there are always conceptual and practical obstacles or restrictions forcing researchers to focus on a particular aspect that can be addressed in practice.
Two important considerations on the choice of a research problem are the nature of the problem and how it has been researched. The nature of the problem is not a difficulty per se, since most topics—even those that may initially seem more qualitative-oriented—can be quantitatively explored by selecting operative concepts that can be converted into measurable variables. Extant research on the issue is crucial in quantitative designs, which typically aim to provide justification for the study by identifying gaps in the literature. Quantitative research is based on a deductive approach articulated in a context of justification or validation, in which theory precedes research.
Gaps may be found in research problems, variables, or even in the type of design. Consider, for instance, the issue of ideology in translation. Ideology is an abstract principle or set of related abstract principles; so, a priori it is not very amenable to quantification. Studies on this topic in TIS have traditionally adopted a qualitative approach based on the exploration of both source (ST) and target text (TT). And yet, there is now work showing that ideology can be approached in a way that allows for quantification of certain variables—e.g., work exploring the impact of participants’ political ideology on their reaction time when translating expressions that may be either congruent or incongruent with their ideological beliefs (Rojo & Ramos 2014). With a close scrutiny, one can always find ways to provide alternative explanations for a problem or complete existing research results. The trick is selecting those variables that are amenable to quantitative exploration.
Posing the research question and picking up measurable variables
Once the research problem has been identified, we can narrow it down and define it more precisely by formulating a clear and simple research question that provides for quantitative measurement. Going back to the previous example on ideology, our research question could be formulated as: how do translators’ political ideology impact their translation process? In posing this research question, the notion of ideology has been narrowed down to the translator’s political view and its effect on translation to the process instead of the product.
But even after the research problem has been more precisely defined in the research question, we still need to operationalize the variables involved, so that they can be measured and quantified. For instance, in most experimental settings researchers observe the effects of independent variables on dependent ones. The former are the ones manipulated by the experimenter and the latter serve to measure the experimental outcome or effect. Apart from defining independent and dependent variables, researchers should also watch other variables that may influence results by intervening or mediating between the independent and dependent variables. These intervening or mediating variables are not measured but, whenever possible, should be controlled by the researcher. In our example on ideology, political ideology would be the independent variable and its effect on the translation process would be the dependent variable. Possible intervening variables could be the participants’ age, certain personality traits, unexpected emotional reactions, or even some anxiety-related factors.
Still, the variables in our selected study are too vaguely defined for measurement and some layers would need to be stripped away to focus on more specific aspects. In quantitative research, at least the dependent variable needs to be measured on a numerical or quantitative scale. For our example on ideology, this restriction entails the need to choose an aspect of the translation process that can be quantified. In the aforementioned work, the dependent variable was reaction time, i.e., the time participants needed to find a translation for the experimental stimuli. Since reaction time is usually measured in milliseconds, it provides a good example of a ratio or numerical scale with a true zero point—0 ms is possible, even if highly unlikely—, in which a given size interval has the same interpretation for the entire scale (e.g., 2000 ms is twice as long as 1000 ms). Other examples of ratio scales are length or number of correct answers on a test.
Besides ratio scales, two other types of quantitative variables are interval and ordinal scales. Like ratio, interval scales are also numerical and have intervals with the same interpretation throughout. Temperature is a typical example of an interval scale, in which the difference between 20 and 40 degrees is assumed to represent the same as that between 40 and 60 degrees. The only difference between ratio and interval scales is that the latter do not have an absolute or true zero point—the zero point in a temperature scale does not indicate an absence of temperature, but rather an arbitrary point on the scale. Typical ratio scales such as reaction time can also be occasionally treated as interval measures when scores are compared against an average. Imagine, for instance, that participants take on average 800 ms to react and that we may want to look at the difference between the observed reaction time and this average level of performance. In such a case, the level of measurement would be made only on an interval scale, since the interval between participants remains meaningful but the ratio element does not—e.g., if a participant scores 1200 ms more than the average while another scores only 200 ms more, the former would still take 1000 ms longer than the latter, but this difference cannot be interpreted as being 6 times as long as the second (i.e., 1200 ms divided by 200 ms) (Fife-Schaw 2012: 30–31).
Ordinal scales are series of values that are ordered, but have no set distance between them. A typical example is that of rank-ordering students according to their abilities into categories such as excellent, good, average, poor, extremely bad. Such a classification tells us that a student may be better or worse than another one, but does not specify how much better or worse. Even if numerical values are assigned to an ordinal scale—as is frequently the case in most Likert-type items, such as 1=strongly disagree, 2=disagree, 3=undecided, 4=agree and 5=strongly agree—, there is still no certainty that the difference between a score of 1 and 2 means the same thing as the difference between a score of 3 and 4. Most psychological test scores measure mental constructs that cannot be directly observed—e.g., attitudes, intentions, personality traits, psychological well-being, etc.—and should, thus, be strictly regarded as ordinal measures, since they infer levels of trait from responses to items about behavioral propensities, but do not measure the construct in any direct way. However, taking a measure to be ordinal has important implications for statistical analysis, since arithmetic means and parametric tests should not be used to analyze them. And yet, researchers have found a way out by treating them as interval values, for instance, assuming that a 5-point difference in an IQ test between someone who scores 75 and someone who gets 80 means the same as the difference between someone who scores 155 and someone with a score of 160. In most research work, this is a common practice allowing researchers to use means and parametric test analyses (Fife-Schaw 2012: 28–29).
In quantitative research, dependent variables need to be countable, whereas independent variables may reflect qualitative or categorical differences. Categorical variables—also known as nominal—are defined by two or more mutually exclusive categories (i.e., one member should not belong to more than one category), but have no intrinsic ordering. Binary categories such as pass/fail and conservative/liberal are typical examples, we could also have more than two categories such as leftist, rightist and centrist politics. The number and type of categories will depend on the research aims and the characteristics of the sample, since we may want to classify participants into conservative vs liberal, or else we may want to add a neutral position. In our ideology example, the independent variable was further operationalized into two opposed political views—i.e., liberal and conservative—which were manipulated to seek congruence or incongruence between the ideological load of the provided stimuli and the translator’s political stance. For the purpose of statistical analysis, numbers may be assigned to observations in each category. For instance, a code value of 1 may be given to conservatives and 2 to liberals. These numbers point to different categories but do not imply more quantity or better quality. Numbers here only serve to provide a reference for computers to calculate means for each category.
Hypotheses, sample size and statistical inference
Once we have selected our variables, we need to formulate predictions on how they are expected to interact, i.e., how we expect the independent variable to influence the dependent one. These predictions are called hypotheses. For statistical purposes, when trying to confirm the validity of a claim there are two basic types of hypotheses: the null hypothesis (H0 or HN) and the alternative one (H1 or HA). The null hypothesis is the assumption that the independent variable has no effect on the dependent one (e.g., stress will have no effect on student interpreters’ performance). Most researchers would not normally pose null hypotheses assuming the absence of effects, but would rather assume that the independent variable will affect the dependent one (e.g., stress will have a negative effect on student interpreters’ performance) and thus formulate alternative hypotheses. However, the null hypothesis is central in testing our predictions, since hypotheses can be falsified or proven untrue, but it is very hard to show they are absolutely true, and always correct. In fact, the alternative hypothesis can only be reached once the H0 has been rejected. Only then can we know that there seems to be some kind of effect and that an alternative hypothesis—hopefully, the one we have posed—may provide a better explanation. This is why the final conclusion should always be given in terms of the null hypothesis. Either H0 is not rejected or is rejected in favor of H1; we should never conclude that the H1 is rejected or even accepted.
Hypotheses can also be classified into correlational or casual, depending on whether variables are predicted to be somehow related or, more specifically, whether a cause-and-effect relationship is assumed between them. Correlations can be positive, when we assume that the relationship between the variables will simultaneously increase or decrease (e.g., the higher the level of student interpreters’ stress, the higher their heart rate during the interpreting task) or negative, when an increase in one variable is associated with a decrease in the other one (e.g., the higher the level of student interpreters’ stress, the lower their score in the task). But correlation does not necessarily imply causation, and should be not used to claim cause-and-effect relationships between two variables.
Hypothetical predictions are always formulated in relation to a population, typically made up of persons or objects (e.g., translations). Since access to all the members of a given population is often impossible, a portion or sample is selected for practical purposes. The assumption is that findings on a given sample will make our inferences valid for the rest of the population, so we need a way to show how confident we are in our inference that what we find in a particular sample also applies to the whole population. And here is where statistics comes in.
The level of confidence can certainly be improved by increasing the size of the sample. The closer its size to that of the total population, the higher the confidence on the validity of findings. Nonetheless, besides practical problems to increase the size of the sample—e.g., lack of professionals willing to participate or having no funding to pay for their participation—there is always a point at which the confidence added by increasing the sample size any further starts to appear less significant. One way to solve these problems is to resort to statistical inference tests, which are to estimate the levels of confidence in accepting or rejecting hypotheses.
When we calculate the likelihood of the null hypothesis, we can incur into two types of mistakes (Fife-Schaw 2012: 24). The first mistake occurs if we wrongly assume that the two variables are related or one affects the other, but in fact they do not. We would be wrongfully rejecting a true null hypothesis. This mistake is called a type I error, and it may occur, for instance, when we end up with an unbalanced distribution of scores for different experimental conditions (e.g., we may have by chance too many students with poor interpreting skills in one group and too many with high skills in the other). In such a situation, the effect on their interpreting performance might not be related to stress but rather to their skill levels. If the probability of this type of error is very low (i.e., lower than 0.05), the null hypothesis can be safely rejected. Note that the probability of making this type of error can never be exactly zero, since samples can always have exceptional cases we are unaware of.
We incur into a type II error when we mistakenly assume that our variables are not related, or that the independent variable does not affect the dependent one, and the opposite is actually true. This type of error may occur, for instance, when too many high skilled students are allocated to the condition assumed to lower the scores (i.e., the stressful condition) and vice versa, viz., too many students with a low skill are allocated to the non-stressful condition. As a result, differences in the scores between the groups might be smaller, and we may wrongly assume that there is no difference when in fact there is.
Most statistical tests provide an estimate of the probability of having made these types of error by calculating the size of the relationships between the observations, the size of the sample and the study design. Accurate estimates of the probabilities of having made these errors improve our degree of confidence on the potential of generalization of data from our sample, and the accuracy of such estimates hinges, in turn, on the use of adequate research tools that provide for precise numerical measurement of our variables.
Selecting precise measurement tools to minimize errors
Successful quantitative research depends on the researcher’s ability to maximize control to minimize errors at every research stage. Apart from inference tests, measurement errors should be accounted for. Measurement errors refer to divergences between the observed value (as recorded) and the true value of a measure. For instance, if we need to measure our participants’ height for an experiment involving heart rate, using centimeters will increase the probability of error, because centimeters are less accurate than millimeters. And the probability of error will be even slightly larger if we use a US tape, since they measure down to 1/16 of an inch, which equals approximately 0.15 centimeters or 1.5 millimeters.
Ratio and interval scales pose no problem to ensure that the points between the scales mean the same in all participants, something imperative for accurate measurement. In these cases, the main drawback may be in the cost of the equipment, since very precise measurement often requires high technology equipment. Most of the equipment used in psychology and TIS research can provide measures of time nearest to milliseconds, a scale that is considered accurate enough in most research areas. This is the case of most reaction time programs (e.g., E-Prime, OpenSesame), eye-trackers (e.g., EyeLink, Tobii Pro) and keyloggers (e.g., Inputlog, Translog-II), which in recent years have been used in translation and interpreting studies to measure participants’ response time, processing speed and cognitive effort in seconds or milliseconds by recording their keypresses or eye movements while typing at a computer keyboard or looking at a screen or at other interlocutors in an interpreting situation (e.g., Rojo, Ramos & Valenzuela 2014, Carl, Bangalore & Schaeffer 2016, Walker & Federici 2018).
A good example of how technology fosters exact measurement can be found in the development from manual categorization of physical expression of emotions—based on Ekman & Friedsen’s (1978) Facial Action Coding System—to computed automated systems that extract the geometrical features of faces in videos and produce temporal profiles of each facial movement. These programs allow researchers to analyze expressions in real time associating them with patterns of basic emotions and producing within seconds or minutes an analysis that could otherwise take hours to identify all the action units—i.e., actions of individual facial muscles or groups of them—and assign them a score in an ordinal scale from minimal to maximal intensity.
Measurements with all these sophisticated technologies are based on ratio scales of time and distance, which guarantee measurement comparability among participants. However, the type of ordinal scales used in most psychological tests and questionnaires raises many more problems. As we saw, labels in Likert scales are probably enough to discriminate between different levels—e.g., between agree and disagree, or even between agree and strongly agree—but cannot ensure that a label means exactly the same for all the participants (e.g., one can never be sure that labels such as “agree” or “disagree” involve exactly the same degree of agreement or disagreement for all participants). Another common problem of Likert scales refers to people’s tendency to choose the more neutral or less discriminating item, a problem that can be overcome by increasing the number of response options available from five to seven or even nine. Nonetheless, increasing the number of options is only advisable up to a certain point. A scale with too many points may certainly seem more precise to the researcher, but also be more confusing for participants, resulting in larger measurement errors (Fife-Schaw 2012: 34).
How to analyze and report quantitative data
When analyzing measurement data, the choice of the correct statistical analysis is also vital to minimize errors. Besides the type of scale or level of measurement used, statistical analysis also depends on the nature of the distribution of scores expected in the population from which the sample is drawn. Depending on the type of distribution expected, we distinguish between parametric or non-parametric tests. Parametric tests make assumptions about the parameters of the population distribution from which the sample is drawn, usually assuming that the population data is normally distributed—i.e., a probability distribution that is symmetrical about the mean and represented by a bell-shaped curve. The shape of a normal distribution is determined by the mean and the standard deviation—i.e., the statistic calculation that shows how close the examples are around the mean in a data set. The steeper the bell curve, the smaller the standard deviation. Conversely, the further the examples are spread apart, the flatter the bell curve and the bigger the standard deviation.
Non-parametric tests can be used for variables that do not show normal distributions, because they are either distribution-free or display a specific distribution in which distribution parameters are unspecified. Non-parametric tests are valid for both normally and non-normally distributed data. And yet, parametric tests are preferred for three main reasons: they have greater prediction power, that is, they may yield answers to questions about our population that are not easily tackled by other means; they have greater statistical power, allowing researchers to detect significant differences when they truly exist; and they allow for more flexible modelling, and accommodate confounding factors using multiple regression.
Parametric tests are generally best applicable to ratio and interval data, whereas non-parametric tests are better for nominal categories, questionnaire responses, rating scores, or assessments on a scale. Still, researchers, lured by the advantages of parametric tests, strive to provide ordinal measures suitable for parametrical treatment. The debate on the adequacy of ordinal measures for parametrical analysis is served, and both sides could have a case. One side argues that many measures lie between ordinal and interval level (Minium, King & Bear 1993), and that the key to the adequacy of ordinal measures is in their quality. If their quality is good, they contend that using parametric tests will probably lead to the same conclusions as using more appropriate non-parametric techniques. Let us go back to the example of a Likert scale. They formally have an ordinal level, but a good Likert scale displays symmetrical categories that facilitate the inference of equidistant attributes and allows it to behave more like an interval-level measurement. This is why, in practice, Likert scales are often viewed as interval.
The other side sustains, in contrast, that the application of higher-level techniques designed for one level of measurement to data of a less sophisticated level often results in nonsensical findings of scarce reliability to draw valid inferences (e.g., Stine 1989). They argue that even if computer test outputs look sensible, such as tables and figures, they will still be farcical, since we have no way of knowing they are leading us to the same conclusions as non-parametric tests. In view of the debate, some scholars (e.g., Blalock 1988) advocate for an intermediate and safer solution that consists in conducting, where possible, analyses on ordinal measures using both parametric and non-parametric tests, confiding on the latter whenever results are contradictory.
Beyond the discussion of statistics experts, the choice of the statistical procedure should also be guided by the conventions to analyze ordinal data in the researcher’s specific field. In TIS, we are used to working with nominal categories and ordinal data that call for non-parametric techniques, although a general preference is also observed to analyze Likert scales by parametric procedures providing for a more complex analysis of different factors. The first steps have already been taken towards more rigorous statistics with specialized work on statistics in TIS (e.g., Mellinger & Hanson 2017). Further work is still necessary, however, to standardize statistical procedures on similar type of data and experimental designs. We cannot agree with Fife-Schaw (2012: 36) when he suggests that only research with serious implications for human life requires a more strict approach to statistical analysis. If we aspire to rigor in TIS research, we should aim at precision and strive for thoroughness and consistency.
Rigor is not only vital for the analysis of results, but also for how such results are reported. Reporting quantitative findings involves three key elements: the statistics carried out, the hypotheses, and the existing theories and studies in the literature. Quantitative findings are first presented in the results section of the paper in numerical form along with the relevant tables, diagrams and figures. Later, in the section devoted to the discussion of results, findings should be reported with reference to the posed hypothesis in order to discuss whether the null hypothesis is rejected and whether, when it is, our hypothesis is the best possible alternative explanation.
The discussion on the likelihood of the alternative hypothesis also requires that we specify how our outcomes compare with existing theories and studies in the literature. In TIS, results need to be collated and examined in comparison with other related studies and associated theories within the field. Finally, the audience addressed by the report is its final judge. Researchers should, therefore, attempt to expound their study, results and interpretations as clearly as possible to their targetted audience so that they can make sense of the whole study and what it means to them and to all the professionals involved—be them teachers, students, translators, interpreters or clients and consumers.
Quantitative research has gone a long way in TIS in the last twenty years, propelled by the momentum of the wave of empirical work from different approaches and areas (e.g., Carl, Bangalore & Schaeffer 2016, Díaz & Szarkowska 2020, Orero, Kruger, Matamala et al. 2019, Vandevoorde, Daems, Defrancq et al. 2020). In recent years, there has been a growing awareness of the need for increasing rigor in research design, the use and application of research tools, or even statistical analysis. So far, advances have been mostly reflected in wishful claims with scarce impact on existing work. A long road still lies ahead of us with regard to many other aspects pivotal for top quantitative research. The guiding principles ensuring the continuity of quantitative research refer to a set of three key ethical requirements that need to be complied with in order to conduct responsible, rigorous and valid research in TIS.
The first requirement imposes the need for TIS research proposals to be of social value, pointing to the contribution of expected results to scientific knowledge, the health or well-being of the populations involved, or the understanding of unresolved problems that may generate proposals for improvement. The originality of the topic and the researcher’s expertise are two preconditions for social value. The social value of research has emerged in recent decades in TIS as an important criterion for novel and responsible research. Well-trodden themes and methods, and single-discipline approaches are either being handled from novel perspectives, or simply giving way to fresh topics (e.g., the study of emotions, the role of personality, or the cultural ecology of translation, the latter being the topic of the IATIS 2021 Conference), innovative methodologies (e.g., physiological indicators, EEG, fMRI, etc.), and interdisciplinary approaches (e.g., psychology, neurology, bilingualism, etc.). Another prerequisite of socially valuable research refers to the need to connect empirical findings and theoretical explanations. The recent surge of empirical research in TIS has resulted in an abundance of empirical findings whose effects are often described without sound theoretical explanations. For TIS research to be of full scientific and social value, there is an urgent need of “converging what and how to know why” (Kotze 2019). Finally, there is also the necessity to beware the perils of tech prone research where technology is put before participants and the ecological validity of the study. It is important to integrate technology into TIS workplace-based research (Ehrensberger-Dow & Massey 2019).
The second requirement refers to scientific validity and rigor, that is, to the need of following every step of the scientific method to ensure methodological rigor and guarantee the production of reliable and applicable results. Scientific validity in quantitative research is supported by many steps, including the correct and bias-free selection of participants, the definition of adequate measurable variables, the provision of valid and reliable measurement using the appropriate instruments, and the compromise to choose and define the statistical procedure used for the analysis of results. Of particular relevance is the process of selecting suitable research tools for a research project. The zest for technological innovation that has characterized TIS research for the last two decades has sometimes resulted in the aimless and excessive use of high-tech equipment that failed to fulfil its initial purpose due to the lax use of the tool and its measurement parameters. Some scholars have already started to lay the foundation for such standards with work on physiological measurement tools (Rojo & Korpal 2020), but more work on measurement standards is necessary to achieve and guarantee this type of scientific validity.
The third requirement points to the need for protecting participants, with a risk/benefit ratio that should ensure minimal risk in relation to potential benefits. There is an obligation to inform participants of the benefits and risks before they actually consent to take part in the study with both the inform consent and the information form. Hiding information on the risks in order to avoid reduction in participation constitutes unethical behavior. Confidentiality of participants’ personal information is also essential to guarantee their protection. Data must be codified to ensure participants’ privacy and anonymity during the whole data collection procedure, when publishing research results, and even at the time of safely storing research documents. While this has been common procedure in areas with a long experimental tradition, such as psychology, in TIS this practice is much more recent, strongly flagged by the codes of ethical conduct and approval processes of universities, research institutes and government organizations. Nonetheless, significant efforts are still needed to harmonize approaches, raise standards, and introduce quality assurance.
New ethical challenges lie ahead related to some of the dangers of our “big data” and “high technology” era. Universities and their research ethics committees need to provide for the perils of open science research for participants’ confidentiality—see, for instance, the scandals related to the OkCupid data release (Kirkegaard & Bjerrekær 2016) and the manipulation of users’ emotions (Kramer, Guillory & Hancock 2014). More published work is needed on possibilities and prospects for technology in TIS research and how we might take necessary measures to reduce the risk of negative impacts. Researchers must learn how to conciliate institutional demands for publication, making data and procedures accessible and transparent, with the moral duty to protect their most valuable asset: research participants.
