Saltar la navegación

Entry

home page help (reference guide) search download (pdf) print broken link

 

 

citation SPA Investigación cualitativa textual

 

contents
Introduction | Qualitative work with texts | Textual Analysis, Discourse Analysis, and Critical Discourse Analysis | Tools for Qualitative Analysis | Research potential

 

Introduction  Introduction 

Research paradigms function as overarching frameworks for interpreting reality and are closely linked to established research methods and techniques (Grotjahn 1987; Williams & Chesterman 2014; Saldanha & O’Brien 2013; Calvo & De la Cova 2024). The term qualitative is inherently polysemous, a feature that often gives rise to conceptual ambiguity and confusions within research contexts.

A first meaning of qualitative refers to studies in which the data under examination are not numerical but predominantly verbal—whether oral, written, or audiovisual (Robson 2016: 5). Accordingly, naturalistic sociological research based on interviews, open‑ended questionnaires, testimonies, and other related instruments is typically classified as qualitative. Hence, research centered on other text‑based materials—such as historical, literary, legal, journalistic, technical, or archival documents—also falls within the qualitative domain. In such texts, relevant data are identified through interpretive procedures that uncover patterns and a wide range of features, including semantic complexity (e.g., intricate expressions or terms, polysemy, cultural references, connotations, double meanings, humor, implicit information); rhetorical devices (such as metaphor, tone, register, or, in interpreting, i.e., strategic silences); ideological markers, perceptions or opinions, communicative intentions, and more. Qualitative analysis is likewise applicable to other forms of data, such as audio recordings—which are typically transcribed and may be enriched with descriptions of non‑verbal communicative cues like pauses, intonation, or speech rate—and audiovisual materials (images, videos, multimodal texts), which are examined by verbalizing the visual elements pertinent to the research aims. Automating qualitative patterns is difficult because their identification and coding generally require some level of interpretation (semantic, contextual, and cultural, functional, lyrical, or aesthetic, etc.) on the part of the researcher. Certain artificial‑intelligence functionalities can assist in identifying and processing particular qualitative aspects that contain explicit textual cues detectable by machines, as in so‑called sentiment analysis (also referred to as emotion analysis or opinion mining). However, qualitative research retains an essential human component and depends on the interpretation of phenomena in scientifically complex, objectified, or operationalized ways—that is, through conceptual descriptions that specify how realities are to be measured or observed—features that would otherwise remain more subjective and reliant solely on the individual researcher’s interpretation. Algorithms provide valuable quantitative data for systematic datasets. However, automation encounters clear limitations when tasked with analyzing texts at a macrotextual level, where they must be approached as complex, hermeneutically interpretable forms of knowledge. Even so, as will be demonstrated, there are already tools capable of supporting and enhancing qualitative analysis processes that rely fundamentally on human judgment.

In its second, more common meaning, the attribute qualitative does not refer to the nature of the data but to the overall research design. It encompasses any empirical study whose feasible method of analysis does not reach—or does not seek to reach—the level of pure statistical inference characteristic of strictly experimental studies, regardless of whether the data processed are numerical or conceptual (qualitative). In this sense, the contrasting attribute quantitative functions as a synonym for experimental. Thus, a study that works with numerical data in non‑experimental, uncontrolled, and statistically non‑representative contexts (for example, questionnaires used in case studies with small samples, or corpus studies with limited scope and low statistical representativeness) would also fall within the qualitative paradigm. This is because its methods, data, and results—whether numerical or conceptual—serve to identify, explore, or describe specific characteristics, phenomena, or observations that arise from examining a particular situation through an evaluation, empirical test, or analysis that lacks the inferential capacity to produce quantitative conclusions potentially generalizable to other comparable situations.

This essential distinction has not always been clearly recognized in translation studies (TS). It is not uncommon to find studies that mistakenly classify themselves as experimental or quantitative simply because they process numerical data or adopt an empirical design that shares certain superficial similarities with experiment-based research, even though their methodology or analytical procedures cannot, from a statistical standpoint, be regarded as inferentially quantitative (i.e., experimental). In many such cases, these works are in fact empirical investigations situated within mixed‑methods paradigms, approaches that are entirely legitimate from a scientific perspective, provided they satisfy the validity and reliability criteria appropriate to their paradigm (Grotjahn 1987).

Diagrama paradigmas ENG
Grotjahn’s Research Paradigms [Graphic interpretation by Morón 2009]

The disciplinary field that has most extensively developed the qualitative paradigm through textual analysis—and its associated techniques and methods—is undoubtedly the social sciences (Strauss & Corbin 1990; Charmaz 2006; Miles, Huberman & Saldaña 2020). During the Cold War aerospace race, in the highly technicist post‑Sputnik era, qualitative and humanistic methodologies—particularly inductive approaches that move from data to theory—did not enjoy the same respect or prestige as deductive (theory‑to‑data), quantitative, or experimental science. Technological competitiveness and the strategic value attributed to scientific progress tended, and in many respects still tend, to be more closely associated with experimental methodologies.

The predominance of purely quantitative approaches over other forms of inquiry remains widespread in many contexts. In response to this perception, sociological researchers—particularly those working inductively—recognized that strengthening the scientific value of their descriptive studies required the consolidation of systematic methodological frameworks. Such systematicity sought to reinforce validity (the extent to which research genuinely measures what it intends to, thereby influencing its potential for generalization) and reliability (the consistency and stability of data, which affect their replicability) (Robson 2016: 15). Consequently, researchers developed comprehensive qualitative research protocols, initially within the school known as Grounded Theory (GT), which—despite its name—is not a theory but a set of inductive textual analysis techniques designed to facilitate the development of new theoretical constructs.

In this sense, the turn toward inductive methodologies in the second half of the twentieth century was, in part, a response to the limitations of the then‑dominant positivist approaches, coupled with a growing interest in explaining complex phenomena that had not been examined in sufficient depth and for which the mere accumulation of large amounts of data failed to yield meaningful insights (Charmaz & Thornberg 2020).

Inductive approaches enable researchers to construct theories both during and after the analytical process. However, the progression from data to theory does not permit the use of preliminary quantification criteria, since what is to be measured is not defined a priori, unlike in conventional quantitative experimental research. That is, this type of data analysis—in our case, analysis of essentially textual data—naturally gravitates toward qualitative and interpretive procedures, either because no established theories exist or because the researcher deliberately sets them aside in order to renew or broaden the perspective on a given topic. Such analysis is necessary, for instance, in text‑based studies of complex or elusive constructs that have not yet been fully consolidated within a field (Calvo & De la Cova 2024), including the key concepts that shape our discipline (error, quality, problem, competence,  equivalence, etc.), as well as in the scrutiny of more intangible textual dimensions that frequently constitute the focus of translation research, such as ideological expression, gendered language, humor, style, tone and register, expressive resources, and other implicit phenomena.

Once inductive methods had been revived and methodologically strengthened, the systematic analytical principles developed in the social sciences over the previous century became reference techniques for other deductive or inductive–deductive research approaches (Miles, Huberman & Saldaña 2020). These principles laid the groundwork for contemporary qualitative, descriptive, exploratory, and other related approaches to working with texts. Their systematicity—rooted in iteration, that is, the cyclical revisiting of data, categories, and emerging theories until greater explanatory coherence is achieved—helped consolidate the view that qualitative studies, alongside more experimental forms of research, also possess scientific validity, even though the notion of validity differs across paradigms. In quantitative contexts, researchers rely on various coefficients of validity and reliability associated with statistical research quality. By contrast, qualitative research is grounded in a concept of validity in which the operationalization of concepts constitutes the core of the study.

Qualitative descriptive research is particularly valuable when no established theories yet exist regarding the reality under investigation, functioning as an initial exploration of a topic. However, such foundational studies may also be significant and scientifically relevant in their own right, as they enable deeper inquiry depending on the nature of the object of study and the objectives of the research.

Empirical research is fundamentally grounded in data derived from real‑world contexts, whether qualitative or quantitative in nature (Grotjahn 1987). At the outset of any study, researchers must make decisions that situate the investigation within an appropriate paradigm: the object of study determines the most suitable types of data, methods, and analytical procedures—a process often referred to as methodological alignment (Hoadley 2004). In essence, the aim is to ensure coherence among the theories, techniques, and forms of analysis employed, and to align them with the characteristics of the data so that the chosen methods are genuinely appropriate for examining the research object. Frequently, a single research object can be approached through both quantitative and qualitative lenses (Grotjahn 1987; Robson 2016). For instance, in textual studies involving metaphors (a qualitative object) or metaphor categories (qualitative data) within a corpus, researchers may quantify their occurrences (a quantitative method) while also interpreting them from pragmatic, stylistic, narratological, or sociolinguistic perspectives (a qualitative method).

Research objects that are easily observable and measurable tend to align with quantitative paradigms, whereas more abstract or insufficiently understood ones—whose variables are not predetermined or stem from novel conceptual perspectives—require explanation and thus call for qualitative, interpretive operationalization (Merriam 1995: 52; Calvo & De la Cova 2024).

To illustrate the existence of multiple research paradigms and how to select the most appropriate one for a given investigation, Grotjahn (1987) proposes examining any study along three dimensions: (1) research design (experimental, quasi‑experimental, or non‑experimental), (2) type of data (quantitative or qualitative), and (3) type of analysis (statistical or interpretive). This three‑dimensional perspective makes it possible to identify not only pure paradigms (quantitative or qualitative) but also six mixed models, thereby expanding the understanding of the study’s methodological nature and facilitating more effective research planning. 

back to top

[heading title] Qualitative work with texts

In TS, as Rojo (2013: 54) observes, research has historically been predominantly qualitative, largely because the discipline’s primary objects of study—text, language, communication, and related phenomena—are inherently interpretive and subjective.

TS has long been characterized by theorization grounded in the researcher’s interpretative capacity, often expressed through abstract philosophical reflections on text, translation, and their associated phenomena. Although such approaches have provided descriptive and explanatory insights that were essential to the discipline’s consolidation, they frequently lack the systematic methodology required to ensure qualitative validity and reliability. Published research often does not articulate methodical, replicable, or consensual techniques, relying instead on individual observation and description. Consequently, there is limited consensus on foundational TS concepts, with multiple labels referring to the same phenomenon and numerous contradictory definitions coexisting for the same term.

 In this context, the limited methodological transfer between qualitative sociological studies and TS is noteworthy, despite the fact that both fundamentally analyze texts—albeit from different perspectives (Calvo & De la Cova 2024). Identifying and operationally defining specific variables makes it possible to process concepts systematically and empirically (Reguant & Martínez Olmo 2014). Such operationalization—understood as the conceptual descriptions that specify how these realities are to be measured or observed—is essential for ensuring greater replicability.

In contrast, quantitative and experimental approaches in TS—such as corpus‑based, cognitive, or terminological studies—draw on more established methodological traditions. This allows researchers to select appropriate research designs from a wide range of options. For instance, cognitive experimental studies may employ techniques such as TAP (Thinking Aloud Protocol), eye‑tracking, or keylogging. In corpus studies, term‑extraction procedures based on frequency counts can be implemented using tools such as AntConc or SketchEngine, and researchers may also rely on official ISO standards for terminological work. Yet despite the prevalence of qualitative research in TS, a paradox persists: these qualitative studies are often the ones that devote the least attention to qualitative methodological operationalization.

In general, the preliminary research design phase—and the methodological reflection it requires—is often overlooked in academic publications, which tend to prioritize the presentation of results (Miles, Huberman & Saldaña 2020: 14). Yet methodological decisions—such as the choice of paradigm, the formulation of clear research questions, the conceptual design, the definition of the corpus, and the selection of data‑collection methods—constitute fundamental analytical acts in the construction of knowledge and should not be neglected in any study. This entry proposes methodological and theoretical axes that may assist qualitative TS researchers in constructing their text‑based qualitative designs and in better understanding the qualitative coordinates relevant to our field.


Qualitative Criteria for Defining a Corpus

Numerous scholars across different disciplines have defined the concept of corpus (Baker 1993; McEnery & Wilson 2001; Tognini Bonelli 2001; Corpas & Seghiri 2007; McEnery & Hardie 2012), a notion that is equally essential in qualitative textual analysis. Broadly speaking, Sinclair (2004) describes a corpus as a collection of representative texts in a given language or linguistic variety, typically in digital format. In descriptive linguistics and TS, the corpus often constitutes the primary object of study, whereas in other disciplines—such as sociology—or in certain research traditions—such as critical discourse analysis—it serves as a means of accessing extratextual realities (e.g., social or ideological contexts). In all these cases, the text remains the cornerstone of research, enabling the use of shared scientific methods and techniques.

Although corpus analysis has historically relied on manual and intuitive techniques (annotations, glosses, index cards, etc.), digitalization and the development of textual‑analysis software have made this work more systematic and efficient. There are now even specialized tools for qualitative corpus analysis, including dedicated software solutions such as CAQDAS.

Formulating an appropriate research question to guide the process provides the rationale for selecting and sampling the corpus and helps ensure that the resulting collection of texts is aligned with the aims of the study (Robson 2016: 249).

It is important to work with “natural” elements, systematically collected in their real contexts for research purposes (Robson 2016; Calvo & De la Cova 2024). Equally essential is the precise definition of the units of analysis that will be employed and subsequently subjected to the coding process (Robson 2016: 249). In TS, these units may take different forms and must be properly operationalized—word, term, token, concept, sentence, segment, unit of meaning, unit of translation, and so on (Calvo & De la Cova 2024: 31).

One of the most debated issues in designing and delimiting the size of a corpus is its representativeness, a concept understood differently in quantitative and qualitative approaches. Quantitative studies emphasize sample size to achieve a degree of generalizability (e.g., number of words, number of documents of the same type, lexical density) and rely on random sampling, which allows researchers to draw statistically or quantitatively meaningful inferences about a given context (Silverman 2005: 126; Calvo & De la Cova 2024: 60). In contrast, qualitative methods typically use non‑random case selection based on explicit criteria such as accessibility, suitability, exemplarity, or uniqueness. In qualitative research, scientific value stems from demonstrating the intrinsic features of the text that make it qualitatively representative of the phenomenon under study (Saldanha & O’Brien 2013: 64–65).

In quantitative sampling, a large sample size is required to achieve greater statistical significance. In contexts that involve qualitative or mixed approaches, however, a smaller corpus may be advantageous, particularly when manual annotation is necessary. With regard to determining an appropriate corpus size, there is no clear consensus even within quantitative research (Tognini‑Bonelli 2001). The quality of the research design outweighs the sheer volume of data (Corpas & Seghiri 2007), a principle that applies even more strongly in qualitative studies. In qualitative corpora, text selection is justified not by randomness but by intrinsic relevance (Saldanha & O’Brien 2013)—that is, by the distinctive features of the texts or by the ways in which they represent the discourse under investigation. It is therefore essential to make explicit the rationale for corpus selection, including its definition and delimitation.

In approaches such as critical discourse analysis or grounded theory, researchers employ theoretical sampling (Mason 1996), whereby texts are selected according to their relevance for answering research questions or developing a theory. This type of sampling is iterative and flexible (Miles, Huberman & Saldaña 2020) and can be constructed and refined as the analysis unfolds; it is therefore not necessarily a design decision made in advance (Silverman 2005). When no new insights emerge during the analysis, so‑called qualitative ssaturation  is reached (Glaser & Strauss 1967; Saldaña 2021), indicating that the qualitative inquiry has achieved sufficient depth (Strauss & Corbin 1998) and that additional data will not yield different results. This criterion ultimately determines the final size of the corpus, since the goal is to understand representative phenomena rather than to exhaust all possible cases. 

Finally, some qualitative studies do rely on closed, predefined corpora, particularly when they involve creative works or unique, extraordinary documents. In such cases, the size of the corpus typically corresponds to the selected work itself, although analyzing it in its entirety is not always required; reaching qualitative saturation is sufficient.


Induction and Deduction for Content Coding

When conducting corpus‑based research, it is necessary from the outset to determine whether the approach will be inductive—developing theory and concepts from the data—deductive—testing pre‑established theories and concepts—or mixed. This debate is closely linked to the distinction between corpus‑driven and corpus‑based studies (McEnery & Hardie 2012: 151; Saldanha & O’Brien 2013: 14).

Deductive coding (or top‑down) is conducted using a set of predefined codes aligned with established theoretical frameworks (Saldanha & O’Brien 2013: 15), which are then applied to the corpus; the researcher already knows a priori what they are looking for. This procedure corresponds to corpus‑based studies, which seek to validate existing theories through corpus analysis (Laviosa 2002; McEnery & Hardie 2012: 6). Inductive coding (bottom‑up), by contrast, generates codes directly from the interpretive process and is characteristic of exploratory and emergent studies: the researcher discovers the nature of the phenomenon in a novel way as the analysis unfolds. This approach is associated with corpus‑driven studies (Tognini Bonelli 2001: 84). 

Saldaña (2021) notes that, in practice, both processes coexist, since inductive analysis produces structures that may later be applied deductively: a theory developed inductively can subsequently be used deductively. As Tognini Bonelli (2001: 85) observes, “there is no such thing as pure induction,” because every observation is shaped by prior linguistic or theoretical intuitions (Sinclair 2004). This methodological complementarity between induction and deduction enables the development of theories more closely attuned to the realities of translation, moving beyond purely academic or preconceived perspectives.


Text function: means or end?

A text can be analyzed from two perspectives: one that treats it as a means of examining social and human phenomena (“discourse as data”), characteristic of the social sciences, and another that regards the text as an end in itself, more typical of linguistic and philological studies. Although both perspectives may employ similar tools, their orientation differs according to how the object of study is conceived. In practice, the boundary between text as means and text as end is not always clear‑cut, and many schools of thought integrate elements of both approaches.

back to top

[heading title] Textual Analysis, Discourse Analysis, and Critical Discourse Analysis

The schools of textual analysis, discourse analysis, and critical discourse analysis encompass a wide array of interdisciplinary approaches whose systematization is inherently complex.

In simplified terms, textual analysis focuses on the text as a linguistic structure (text‑analytic approach), examining its form, syntactic organization, and cohesion, whereas discourse analysis concentrates on language use in its social context, exploring how discourse reflects and shapes social relations, ideologies, and power dynamics. In practice, however, these two levels are often closely intertwined. 

Critical discourse analysis, conceived as a critical branch of discourse analysis, highlights the social implications of language, understood not only as a means of representing reality but also as a tool that actively contributes to constructing it. This approach has been applied in TS to examine issues of ideology, power, and representation in translated texts, particularly in domains such as the media, immigration, gender, and censorship—issues that concern “how groups of people conceptualise themselves, their social setting, other groups of people and the issues that matter to them” (McEnery & Wilson 2001: 133).

In general methodological terms, studies grounded in textual analysis and discourse analysis tend to rely on more structured, deductively oriented theoretical frameworks, whereas critical discourse analysis favors more flexible and inductive methodologies aimed at the critical examination of social phenomena through language. The choice of techniques and methodological positioning must be aligned with the research design, the nature of the object or data under study, and the type of analysis being undertaken.

School/ Technique

Design (method)

Data (object of study)

Analysis

TA

Deductively, intuitively identifying textual patterns following theories. Technique is not described in detail.

Text and its elements (syntax, morphology, etc.). Text as an end.

Descriptive, qualitative

DA

Method is not described in detail.

Discourse. Text as a means and as an end.

Qualitative. Interpretation standpoint is described in detail (contextual, pragmalinguistic dimension).

CDA

Rarely illustrated.

Object of study is thoroughly defined, constructed. Mostly text as a means.

Critical and ideological coordinates are established for data interpretation.

Quantitative CL

Method of analysis is well-established (e.g. quantitative software).

Quantifiable text elements in corpus. Text as a means and as an end.

Different forms of interpretation are possible (statistical, interpretative, etc.).

Qualitative CL

Rarely described

Interpretative elements in corpus. Text as a means and as an end.

Form of interpretation is expressly indicated.

CA

Greater effort to highlight and validate the method of analysis used (GT, QDA, etc.). Inductive-deductive.

First cycle coding, second cycle coding.

Datum as a medium + contextual-social dimension. Datum as a medium.

 

Iterative coding techniques to interpreting the data sociologically, ideologically, etc. Standpoint is usually well described.

GT

Pure inductive coding. Constant comparative method: open coding, axial coding and selective coding:

Datum as a medium.

A grounded theory inductively derived from corpus/data.

Design–Data–Analysis Scheme in Different Research Schools

[Calvo & De la Cova 2024]

Corpus Linguistics

Corpus linguistics (CL) was initially applied to collections of monolingual texts, with a particularly significant impact in fields such as lexicography and terminology. Since the 1990s, however, it has been adapted to TS—especially following the contributions of Baker (1993, 1995)—giving rise to corpus‑based TS (Baker 1993; Laviosa 2002; Corpas & Seghiri 2007, among others). Applied to translation, CL enables researchers to observe the object of study directly and has played a key role in the development of the descriptive branch of TS. Although its orientation is predominantly quantitative—drawing on metrics such as word length, type–token ratio, or statistical significance—many scholars advocate combining it with qualitative approaches that contextualize and explain the phenomena observed in greater depth. This mixed approach is particularly valuable for analysing factors such as genre, textual purpose, or translation orientation. Authors such as Mason (2009) and Olohan (2021) emphasize the need to move beyond statistical data and also examine context, communicative intention, and discursive function—elements that are fundamentally qualitative—in order to achieve a more comprehensive understanding of the translation process.


Content Analysis

Content analysis is a methodology used to examine textual documents with the aim of understanding a social or communicative reality. It emerged in the early twentieth century in the United States as a quantitative technique for studying the press and other mass‑media outlets (Robson 2016: 351; Neuendorf 2017), but over time it has evolved toward more interpretive and qualitative approaches, incorporating techniques such as interview analysis, open‑ended surveys, and direct observation (Robson 2016).

Although it shares certain objectives with methodologies such as discourse analysis, content analysis places greater emphasis on the text as a means of understanding social contexts, whereas discourse analysis focuses more on linguistic styles and strategies (Robson 2016: 371). Content analysis also frequently incorporates computer‑assisted qualitative analysis through CAQDAS tools (Computer Assisted Qualitative Data Analysis Software).


Grounded Theory

Grounded Theory (GT), developed by Glaser and Strauss (1967), is not a theory in itself but an inductive methodology for qualitative analysis whose purpose is to construct concepts or theories directly from the data—that is, grounded in empirical evidence (Robson 2016: 161). According to Strauss and Corbin (1990: 23):

A grounded theory is one that is inductively derived from the study of the phenomenon it represents. That is, it is discovered, developed, and provisionally verified through systematic data collection and analysis of data pertaining to that phenomenon. Therefore, data collection, analysis, and theory stand in reciprocal relationship with each other. One does not begin with a theory, then prove it. Rather, one begins with an area of study and what is relevant to that area is allowed to emerge.

A bottom‑up approach based on a cumulative and iterative coding process, GT is particularly useful when there is limited prior literature on the phenomenon under investigation (Saldaña 2021: 72). GT aims to generate theory directly from the data, without imposing theoretical categories from the outset. Although GT and content analysis both employ coding and categorization of qualitative data, their objectives differ. GT seeks to develop theory from the data through a systematic inductive process, whereas content analysis identifies categories to address predefined research questions and may follow an inductive, deductive, or mixed approach (Cho & Lee 2014: 5).


Schools and Coding Techniques

Coding is the fundamental process in all qualitative analysis. According to Strauss and Corbin (1990: 57), it consists of “the operations by which data are broken down, conceptualised, and put back together in new ways.” It involves interpreting, categorizing, and explaining the data in detail (Böhm 2004: 270–271; Silverman 2011: 70). When carried out iteratively and rigorously, coding enables the theoretical validation of qualitative analysis, making it more replicable, transparent, and adaptable. Coding therefore entails identifying passages that exemplify a shared theoretical or descriptive idea and grouping them under a label or “code” that represents that concept (Gibbs 2007: 38). 

This process goes beyond superficial labelling or marking (labelling, indexing, or tagging), as it requires continuously comparing and relating coded units to one another and to established concepts—what is known as the constant comparative methods (Charmaz 2006). It alternates between the inductive (emergent) reading of the data and the incorporation of existing theories or models (deduction) (Srivastava & Hopwood 2009: 77). A code is a word or short phrase that, according to Saldaña (2021: 3), “symbolically assigns a summative, salient, essence capturing, and/or evocative attribute for a portion of language based or visual data.” It is a construct generated by the researcher, who interprets and symbolizes the meaning of a datum to facilitate subsequent processes such as categorization, formulation of assertions, or theory building. Although some consider coding merely a preliminary phase before the “real” analysis, others—such as Miles, Huberman and Saldaña (2020: 63)—regard it as the very core of qualitative analysis: “coding is analysis.” 

Análisis y codificación de datos (Saldaña 2021)
Data Analysis and Coding [Saldaña 2021]

These proposals from the social sciences have given rise to methodological frameworks that are both useful and applicable to TS. In particular, the distinction between first‑cycle and second‑cycle coding methods (Saldaña 2021) is especially significant.

Coding techniques were initially consolidated through Grounded Theory (GT) in the twentieth century. From this inductivist perspective, they were progressively transferred to what later became known as Qualitative Data Analysis (QDA), which may follow a deductive or an inductive–deductive method. In GT‑specific coding, identifying the core category—that is, the theory—emerging from a dataset involves a structured analysis carried out through three main iterative coding stages: open coding, axial coding, and selective coding (Strauss & Corbin 1990: 58; Böhm 2004: 270; Robson 2016: 463).

  • Open Coding: This stage consists of breaking down, examining, comparing, conceptualizing, and categorizing the data (Strauss & Corbin 1990: 61). The content is analysed exhaustively, relevant excerpts are identified, and initial codes or conceptual categories are assigned. It is generally advisable to begin by coding a representative portion of the corpus and then extend the coding scheme to larger sections (Robson 2016: 163, 463; Böhm 2004: 271).
  • Axial coding: “a set of procedures whereby data are put back together in new ways after open coding, by making connections between categories. This is done by utilizing a coding paradigm involving conditions, context, action/interaction strategies and consequences” (Strauss & Corbin 1990: 96). During axial coding, the working data consist of the preliminary codes generated during open coding. At this stage, relationships among these codes are identified and articulated. Axial coding is crucial for theory generation, as it enables the researcher to link categories and concepts at a higher level of abstraction (Robson 2016: 463; Böhm 2004: 271).
  • Selective coding: “the process of selecting the core category, systematically relating it to other categories, validating those relationships, and filling in categories that need further refinement and development” (Strauss & Corbin 1990: 116). Axial coding provides the framework for abstracting a core category, which functions as the integrative nucleus of the analysis. This category may be refined as the study progresses. Identifying a category sufficiently abstract and general to serve as a clear theoretical axis is not always straightforward (Robson 2016: 463; Böhm 2004: 273).

During this process, keeping memos (analytic conceptual notes) is essential, as they document the rationale behind potential codes and later support the reasoned revision of initial coding proposals, enabling these to evolve into more stable categories.

Since GT is an inductive methodology, it is advisable to create original category labels, avoiding the reuse of already established ones as far as possible (Silverman 2011: 70). Because it is a purely inductive approach, the aim is to develop new theories based on a little‑studied dataset or, in some cases, to approach a reality from a new perspective or with a new methodology that purposefully and justifiably departs from pre‑existing theories.

Both QDA and GT apply the constant compative method to conduct systematic analysis and foster theory generation. This method comprises four stages: (1) comparing cases associated with each category, (2) integrating categories and their properties, (3) delimiting the theory, and (4) writing the theory (Glaser & Strauss 1967/1999: 105).

One of the most frequent criticisms of GT, according to Robson and McCartan (2016: 162), is that it is unrealistic to conduct research without reviewing theoretical frameworks, prior ideas, or assumptions—as if the researcher were a tabula rasa (Charmaz 2006). Glaser (2004) argues that although GT must remain essentially inductive, it cannot ignore pre‑existing concepts and theories; these should be incorporated into comparative analysis once the core category has been identified and conceptual development is underway (Glaser 2004: 12). The issue is not whether existing theory should be ignored, but when it should be integrated: the traditional deductive hierarchy (theory → practice) dissolves, and the theoretical framework is incorporated when the data require it, even after the main analysis phase. The strict application of pure GT techniques is complex. In practice, the boundary between inductive and deductive approaches is not clear‑cut. An inductive design follows a different logic from one that presupposes finding already described phenomena in the data, yet it may acquire more deductive traits as inductively derived concepts take shape. Conversely, a deductive design may adopt inductive features when existing categories fail to adequately capture the phenomenon under study. A rigorous inductive–deductive study must therefore distinguish genuinely inductive elements (those emerging anew from the data), deductively applied concepts, and those revised or reinterpreted inductively in light of the data.

Although GT has not been widely exploited in TS, its methodological potential has long been acknowledged. For example, Hubscher‑Davidson (2011: 6) was among the first authors to explicitly highlight the relevance of GT for qualitative research in TS.

GT has been applied primarily in studies analysing data gathered through sociological or ethnographic techniques—especially interviews, open‑ended questionnaires, or testimonies—as in Olohan (2014), Pöllabauer (2012), Rico & González (2022), or Sun (2011). The limited qualitative methodological tradition in TS means that many studies referencing GT actually operate along an inductive–deductive methodological axis rather than within a strictly inductive approach. A representative example of this hybrid methodology is the work of Désilets, Melançon, and Patenaud (2008).

Beyond its application in interview‑based and sociological data studies, GT also provides a fertile methodological framework for textual analysis in TS, both for source texts and translations, particularly in research exploring translation strategies, errors, reformulation patterns, or poorly defined theoretical dimensions. Relevant contributions in this line include De la Cova (2017), Szymylsik (2019), and Calvo & Morón (2020), who apply GT’s iterative logic to textual analysis. Particularly noteworthy is Wehrmeyer (2014), who develops a reception‑based model for TS by comparing source and target texts through audience expectation norms.

Ultimately, although distinguishing between deductive and inductive categories is essential to avoid methodological confusion, GT remains a valuable epistemological framework for investigating ambiguous concepts or exploring new domains in TS. Its constant and iterative coding logic enables the development of robust categories grounded in the data without relinquishing critical engagement with theory.

In practice, many GT‑inspired approaches function as genuinely inductive–deductive methodologies. What often attracts translation scholars is the richness of GT’s systematic constant comparison process rather than strict inductive purity. The limited transfer of qualitative sociological methodologies into TS, combined with the scarcity of systematic qualitative studies in translation that explicitly describe their coding procedures, has hindered the broader adoption of this inductive methodology.

Thus, when a research process does not follow a purely inductive logic but incorporates pre‑existing theoretical categories deductively from the outset, it is more appropriate to refer to Qualitative Data Analysis (QDA) rather than strictly to GT, as the two are not methodologically equivalent: the role of theory differs fundamentally in each design.

QDA adapts the techniques developed for GT to inductive–deductive and deductive processes, and it likewise conceptualizes coding as a cyclical procedure. The first cycle, known as Initial Coding or First Cycle Coding (Saldaña 2021), may be considered analogous to open coding in GT. At this stage, the researcher assigns preliminary codes to units of data, but the process is reflexive and does not rely on a single reading: the researcher returns to the corpus as many times as necessary to connect emerging ideas. The outcome consists of one or several initial codings accompanied by annotations, comments, and memos.

In the second coding cycle, the focus shifts from the text itself to the preliminary codes generated in the first phase. As Saldaña (2021: 89) notes, this cycle demands “analytic skills such as classifying, prioritizing, integrating, synthesizing, abstracting, conceptualizing, and theory building.” The task involves identifying patterns that may relate to phenomena, causes, types of relationships (opposition, inclusion, dependency), the actors involved, and their interconnections. By grouping and linking codes, the researcher constructs a higher‑order, coherent analytic system—comprising hierarchies and conceptual maps based on criteria of relation, inclusion, and exclusion—which constitutes qualitative theorization.

Robson (2016) describe this theorization process as thematic coding and define five basic steps:

  1. Familiarization with the data (reading, transcription, preliminary ideas).
  2. Generation of initial codes (with or without a prior conceptual framework).
  3. Identification of themes (grouping codes, reviewing coherence).
  4. Construction of thematic networks.
  5. Integration and interpretation (use of matrices, schemas, interpretation of patterns).

During the analytical phases, the process shifts from observation to operationalization: “clear operational definitions are indispensable so that they can be applied consistently by a single researcher over time and multiple researchers will be thinking about the same phenomena as they code” (Miles, Huberman & Saldaña 2020: 77). Such operational definitions enable the validation (or refutation) of studies in simultaneous or subsequent research.

Saldaña’s (2021) comprehensive work on QDA presents a wide range of techniques, perspectives, and coding methods, along with numerous concrete examples of QDA research (from a general sociological rather than a specifically translation‑oriented perspective, though highly valuable for understanding the process). In the following image reproduced from Saldaña (2021), one can observe how text—understood as the primary datum (open coding in GT or First Cycle Coding in QDA)—is transformed into codes as a higher‑level data type, then crystallizes into categories (axial coding in GT or Second Cycle Coding in QDA), and ultimately solidifies into theory in the form of themes and concepts (selective coding). These cycles should not be interpreted as linear; rather, as noted earlier, they constitute an iterative back‑and‑forth process that continues until coding is consolidated and categories reach qualitative saturation.

The final set of codes is usually represented in a codebook (coding book or coding scheme). It is advisable to keep successive versions of this codebook throughout its development in order to maintain what is known as a living codebook, a research tool in its own right that helps document the evolution of the analysis and includes preliminary codes, definitions, examples, and justifications for decision‑making. This record facilitates the traceability of the analysis, enhances reliability, and supports intersubjective consensus in team‑based research or in future deductive investigations (Reyes, Bogumil & Welch 2021).

back to top

[heading title] Tools for Qualitative Analysis

Aplicación cualitativa de códigos para detectar problemas de traducción en Atlas.ti Web
Example of a qualitative application of codes to identify translation problems in a corpus of English‑language judicial decision, Atlas.ti Web interface [Ongoing research conducted by E. Calvo]

In qualitative data analysis (QDA) studies, a wide range of manual methods have traditionally been used for marking and annotating content, including marginal notes, comments, glosses, memos, flags or markers, underlining, analysis cards, tables and templates for organising information, and folders, among others (Robson 2016: 464; Saldanha & O’Brien 2013: 231). Many of these procedures require prior preparation; for instance, when coding manually with digital tools, pre‑editing the texts is often recommended to facilitate subsequent processing. However, such approaches present notable limitations, including restricted dynamism in data retrieval and reduced efficiency—particularly in contrastive studies that require the analysis of aligned bitexts. These constraints underscore the need for more specialised and powerful tools for qualitative research in TS.

According to Fantinuoli (2016), the quantitative tools available to date were originally designed for pedagogical purposes or for monolingual applications in linguistics and lexicography (WordSmith Tools, AntConc, TextSTAT). These tools are agile and efficient, enabling the automatic processing of complex file formats, the use of corpora without prior preparation, and even the creation of corpora directly from the web (Sketch Engine). However, they still present limitations in professional translation-related workflows and in their application to TS, particularly when the object, method, or outcomes of the study are qualitative, since machines cannot interpret texts adequately at that level (Fantinuoli 2016). From an analytical standpoint, these quantitative tools do not generate research-ready representations of results comparable to those produced by some CAQDAS tool; here the researcher must largely construct both the findings and their visual or conceptual representation.

In this context, several digital tools have been used or adapted for TS:

  • Aligners (Trados, MemoQ, LF Aligner): useful for working with sentence‑level segments and bitexts, although they do not incorporate advanced annotation or result‑representation functions.
  • General tools, such as OCR or transcription software (ExpressScribe), which facilitate pre‑editing and the conversion of non‑textual data into processable corpora.
  • In the absence of specific tools, some researchers have developed or adapted annotation or coding software such as EXMARaLDA (Extensible Markup Language for Discourse Annotation), applied to the TiPP project at the Autonomous University of Barcelona (Orozco 2018)
  • Finally, CAQDAS software (Computer Assisted Qualitative Data Analysis Software) offers the most comprehensive functionalities for qualitative analysis in TS.

CAQDAS tools such as NVivo, ATLAS.ti, MAXQDA, Dedoose or Delve enable researchers to code, classify, and visually represent data—through concept maps, word clouds, or charts—and to develop theories from emerging coding structures. Although their use in TS remains limited, their potential is clear, particularly for monolingual textual analysis of original or translated corpora.

Finally, although some studies in TS have already employed CAQDAS—for instance, in ethnographic research, discourse analysis, or literary studies—their application to qualitative textual analysis from a translation‑oriented perspective remains incipient. This is due in large part to the limitations these programs still present for contrastive studies, as they are not designed to manage aligned multilingual corpora. While it is possible to work with bilingual tables or bitexts, and creating hyperlinks between source and target texts can be particularly helpful, the ideal scenario would be the development of a CAQDAS tool specifically designed for bitexts.

back to top

 Research potential Research potential

The range of study techniques inherent to, or compatible with, qualitative methods of textual analysis—whether the text serves as a means of understanding a broader research object or constitutes the object itself—offers multiple avenues for addressing key theoretical challenges within TS.

The main phenomena examined in translation often involve complex and subjective concepts—such as translation problem, translation competence, translation error, quality, and a broad array of textual, stylistic, discursive, and functional notions. In a disciplinary context where a single object of study may be described with multiple labels, and a single label may encompass a wide spectrum of distinct realities, the most effective research response lies in precise conceptual definition and in systematic methodological techniques that enhance qualitative validity and reliability.

In both deductive studies and inductive proposals, the systematicity of the methodologies applied and the conceptual validity they sustain are what will position TS at a higher theoretical level capable of fostering disciplinary advancement. This methodological and theoretical robustness is particularly crucial today, given the disruption introduced by artificial intelligence and its data‑driven approaches to textual processing. In contrast to this predominantly quantitative paradigm, qualitative studies reaffirm the centrality of hermeneutic interpretation, interpretive analysis, and a deeper understanding of the translatable or translated text. This interpretive orientation also most clearly exposes the limits of automated processing and underscores the need to support—or triangulate—purely quantitative investigations with systematic qualitative research.

Texts constitute the primary research material in TS and lie at the intersection of language and culture. To date, no attempt at quantitatively systematising language has achieved comprehensive coverage. A qualitative approach, by contrast, links the notion of quality to contextual constraints and expectations, as well as to the intentional function that shapes textual complexity, thereby aligning with the logic of major functionalist authors in TS. Systematic qualitative coding of texts—understood, depending on the case, either as the object of investigation or as the means through which it is conducted, whether following inductive, deductive, or inductive–deductive principles, and adopting a top‑down or bottom‑up perspective in line with iterative QDA techniques—is suitable for a wide range of translation‑oriented studies, including,

  • Contrastive studies, aimed at coding phenomena identified in the target text for purposes such as quality assessment, the identification of strategies employed, and other related aspects.
  • Studies on the source text and its nature as the beginning of the translation process—for example, to detect translation problems or other textual phenomena that may not be apparent in the target text (Calvo & De la Cova 2024).
  • Studies on ideological aspects or discursive markers that enable the analysis of social contexts (censorship, gender, discrimination, political views, etc.)
  • Studies on argumentative or narrative quality (soundness of arguments, originality, writing quality, textual cohesion and coherence, fluency, eloquence, false or distorted information, plagiarism, cognitive dissonance, etc.)
  • Studies on style, tone, voice, register, and other literary or expressive qualities.
  • Studies on intangible or implicit aspects of discourse, such as humor, irony, sarcasm, double meanings, connotations, profanity or coarse language, learned expressions, metaphors, and other literary devices.
  • Studies on complex subordinate contexts, such as visual, audiovisual, or multimodal texts, orality‑based discourse, rhetorical devices, etc.
  • Studies with a strong cultural, local, legal, historical, or social component, requiring extratextual interpretation.

back to top

Creado con eXeLearning (Ventana nueva)