
Add us as your preferred source on Google to see more of our content in Search.
This article provides a concise literature review on how evaluation is defined in the development context and how consultants can represent their evaluation experiences in a more effective way. The author, who has a sociological background, has also conducted various evaluations and faced several dilemmas due to misunderstandings in identifying her own evaluation experiences. Research and evaluation are distinct, and job experiences should be selected accordingly. Evaluation has become a complex sector, with unique approaches and terminology that should be well understood.
Introduction: the weight of evaluation in the development sector
DevelopmentAid platform registered 106,875 job posts on its platform in 2025, of which 25,996 were related to two correlating sectors: Project Management and Monitoring & Evaluation (taking the first place in the ranking of the 10 top hiring sectors). Organizations were looking for research analysts, monitoring and evaluation specialists, and quality assurance experts.
In addition, 1,525 tenders were available in the evaluation sector for individual consultants in the last year, prioritising processes like ongoing oversight of a project’s progress against its planned results, spotting risks, gaps in implementation, working on timely adjustments to maintain the pace of programs’ development, midterm reviews to evaluate progress, final evaluations, evaluating relevance, effectiveness, impact, and sustainability. Consultant profiles sought were calling for impact assessment, civil society experts, data analysts, and experts in M&E plan development and implementation. Some of the most active funders included the International Labour Organization, the World Bank, and UNESCO.
These analyses reflect the wide range of evaluation activities as well as of the required expertise, of which a civil society background seems to be highly relevant, since project management and evaluation skills often go together with the voluntary or non-governmental sector.
Definition of evaluation
Evaluation has a long history, but evaluation theories, as we know them now, first appeared in the United States in the 1960s, with the establishment of Great Society programs that needed evaluation. Since then, as an impact of the largest international aid organisations (like the OECD, the UN, the World Bank), evaluation has been gaining importance and going through professionalisation. The most widely known evaluation system comes from the OECD and its Development Assistance Committee (DAC) from 2002, which defined evaluation as:
“Evaluation is the systematic and objective assessment of an ongoing or completed project, programme or policy, its design, implementation and results. The aim is to determine the relevance and fulfilment of objectives, development efficiency and effectiveness, impact and sustainability”.
The main purpose was to increase the effectiveness of international development programs and to perform robust, informed, and independent evaluations. The uniqueness of the OECD-DAC system is the five evaluation criteria: relevance, effectiveness, efficiency, impact, and sustainability. The OECD-DAC definition and evaluation criteria became an internationally agreed reference point for other international organisations and the European Union.
Evaluation approaches in the development sector
The standardisation of evaluation resulted in evaluation policies of the international organisations, describing concept, purpose, principles, rules and responsibilities, measures, and dissemination. The Terms of Reference (ToR) that consultants meet first reflect the evaluation policy of the clients.
The United Nations (UN) founded its “Inter-Agency Working Group on Evaluation” in 1984 (since 2003 it is called UNEG) to introduce M&E systems in operational activities. Evaluation norms include utility, credibility, independence, impartiality, professionalism, transparency, and use of learning.
The European Union accepted the definition of the OECD and formulated its evaluation methodology for its external assistance around the 2000s. In the light of good governance and the “evaluate first” principle (EC, 2013), the EU is engaged in a strong culture of accountability and learning. Learning refers to generating evidence-based knowledge about what works and what does not, and under what conditions. The EU proposal system also involves “in-built” evaluation mechanisms. In terms of international co-operations, work package structure includes an evaluation and/or a quality assurance component, while in national proposals the Log Frame matrix and the Gantt chart serve the same purpose, at the same time disseminating the culture of evaluation among actors.
The World Bank also follows the international evaluation principles, with a focus on maximising impact.
As a recent development, most aid organisations evaluate the evaluation reports and provide open access to them with relevant information on the subjects of the evaluation (project/program goals, implementation, final evaluation report and its acceptance). UNDP’s Evaluation Resource Centre can be mentioned as a sample here.
By today, policy frameworks of these influential international organisations involve equity considerations and recommendations as well.
Finally, it is important to note that these evaluation regimes usually have a separate quality assurance section as another layer of evaluation. Quality assurance refers to the definition of “high-quality” (of data, products, reports), and works as an analytical tool, in contrast to quality assessment, which is an alternative evaluation tool. For instance, in the European context, quality assurance refers to the clarity of the whole evaluation process. The OECD and the UN also employ quality standards for improving the quality of evaluations and to minimise failures (e.g. by comparisons, ratings and rankings).
Evaluation models
Evaluation methodology can be derived mainly from sociology and, in a smaller part, from market research. Despite similarities in methodologies (quantitative, qualitative) and ethics (independence, objectivity), research and evaluation are distinct due to the different approaches. While sociological research is descriptive, evaluation is prescriptive (measures results against predefined standards) in its nature, and they use different terminology.
Summative and formative evaluation
Summative and formative evaluation give the traditional approach in evaluation, and many evaluation tools are rooted in this approach, and they are widely used in many fields. The terms “summative” and “formative” evaluation was first introduced by Michael Scriven, a philosopher of science, who mentioned them in his book “The Methodology of Evaluation” in 1967. With simple words, formative evaluation is conducted during a process to improve quality, while summative evaluation occurs at the end to define overall achievements. These are complementary methods, with different characteristics, objectives, and implications (see Table 1).
Table 1: Differences between formative and summative evaluation
Monitoring, M&E, MEL, MEAL and MERL evaluation models
Monitoring and evaluation often cover a complex formative and summative evaluation; however, in some cases, clients ask for a specific evaluation, like MEL, MEAL or MERL (see Figure 1):
Figure 1: Monitoring and Evaluation-based evaluation models
- Monitoring covers a routine (formative, iterative) assessment of resources, activities, and results.
- Evaluation is a periodic (inception, interim, midterm, final) assessment of an ongoing or completed program. It highlights intended and unintended consequences.
- Learning uses the information generated from M&E intentionally to continuously improve performance.
- Accountability puts a greater focus on communication, transparency, and dissemination in management responsibilities.
- Research adds an exploratory element to evaluation, with the questions: what works and under what conditions, or what happens if…?
In practical use, monitoring provides a formative assessment of a project or a program. Evaluation often completes monitoring with a summative approach (done by the same evaluator). If not, the evaluator performs an overall and independent summative evaluation. Learning and research both give some innovative elements in the evaluation process, where the evaluation itself can bring new aspects to improve program or project performance through its results and feedback loops.
Research is specific in a sense that it refers to a research component (in international projects, a “built-in” pilot test and period, with no predefined results), which is followed through an impact assessment (using additional, specific formative and summative, i.e., monitoring and evaluation tools)
These models, from monitoring through evaluation to research, are catching the increasing complexity of projects and programs (e.g. the multinational context, incorporated pilot research, interdisciplinary nature, etc.).
SWOT, PEST and PESTEL (or PESTLE)
SWOT is widely spread in project management and evaluation; however, PEST and PESTEL are rooted in market research and have a smaller impact on evaluation models. Both are used to assess conditions and environments for strategic planning, but also for evaluation. While SWOT has a dual focus and a more introspective orientation, PEST and its alternative, PESTEL are explicitly focusing on environments:
Table 2: Comparison of SWOT, PEST and PESTEL (PESTLE)
The purpose of SWOT is to identify positive and negative aspects affecting progress, both internally and externally. PEST and PESTEL are used when SWOT fails, and a macro-scale analysis is required.
Hence, PEST focuses on the environment of interventions, so it assesses:
- political (governments and policy background, monetary and fiscal policies, reforms of labour market and workforce, regulations, political stability, initiatives),
- economic (inflation, interest rates, economic stability and growth, demand/supply trends, budgets, operational costs, exchange rates and pricing strategies, taxes)
- and social circumstances (trends in population, demographics, workforce trends and domestic unemployment, trends of work, situation of a sector, cultural trends and awareness of issues).
The PESTEL is more distinctive in defining influential factors, and is adding:
- technological trends (infrastructures, advancements, innovations, emerging technologies, actors).
- environmental aspects (natural disasters, effects of climate change, sustainable practices).
- and legal issues (regulatory changes and stability, laws, taxes) to the assessment.
Probably, PEST is also capable of measuring the environment, so it depends on clients’ terminology (PEST or PESTEL).
Quality Assessment (QA)
In large-scale European Commission projects, quality assessment is a fundamental component, since it helps handle complexity and minimise potential failures. Its aim is to ensure that the objectives and requirements are met at every stage of the project life-cycle, and in a high-quality. Quality assessment highlights challenges and allows risk mitigation, and improves team morale and engagement. It provides timeliness and integrity at the project and management level. In terms of methodology, quality assessment might employ risk assessment (SWOT), iterative protocols (participatory tools like internal surveys, meetings, observation), output (peer)reviews, timetables (e.g. a Gantt chart); tools enabling progress follow-up and monitoring. Lately, quality assessment incorporates data management, too (quality data collection, sensitive data, handling data), due to GDPR rules.
Evaluation ethics
We need to mention ethical considerations in evaluations in a few words. International aid organisations define their codes of conduct, including principles like high-quality, credibility, strong accountability, evidence-based information, accurate and reliable methodology, learning, and utility.
Values concerning evaluators are honesty, expertise, and independence. These all enable evaluations to be a good base for further policy-making. “From knowledge to policy” is a significant shift that can be seen in the evolution of evaluation theories and models.
Evaluability and the inception phase
Finally, regarding the evaluability of projects and programmes, an ethical note could be brought here, namely the question of the inception phase, which used to be part of the Terms of Reference. The inception report is the initial stage in an evaluation process, based on the analysis of the context and available documentation. It is a kind of research plan that defines the evaluation concept and methodology.
Writing an inception report is quite time-consuming. Unfortunately, recent practice shows that development organisations ask for the inception report (based on some shared documents attached to the ToR) besides consultants’ CVs. This practice generates uncertainties, since costs of such freely done inception reports are not calculated in the evaluation process once the consultant wins a tender. In case of not winning applications, consultants just did the job.
Summary
All the above-mentioned evaluation models can be placed on the formative-summative scale, and we can see their combinations in evaluation assignments.
Figure 2: Formative and summative evaluation models
Figure 2 shows formative and summative tools for evaluation. Hence, individual consultants can better identify their own job experiences by being aware of these tools.
It helps distinct their research and evaluation job experiences, too, where they seem to overlap.
ToRs usually predefine evaluation methodology; though, there are cases where evaluators choose the best-fitting methods. Some theories (e.g., the Theory of Change) are different from the traditional formative-summative approach, since it does not have predefined standards, and methods are selected accordingly to the occurring changes.
Key takeaways from this article
This article overviewed evaluation models applied through the development sector and, to some extent, the evolution of evaluation. It highlighted that research and evaluation are distinct activities despite the similarities in methods and ethical considerations. It recommended practices for individual consultants to identify and select evaluation job experiences along with evaluation models, and to tailor CVs with greater efficiency for evaluation assignments. Finally, there showed up some unfavourable trends in evaluation jobs.