Every archaeological interpretation begins with evidence.
But every archaeological interpretation also begins with assumptions.
That realisation stayed with me throughout the years I spent documenting Palaeolithic artefact occurrences across the Wainganga–Wardha Basin in central India. As more sites appeared on my maps, I expected the patterns to become clearer. Instead, a different question kept becoming more difficult to ignore.
How much of the organisation we recognise in archaeological landscapes genuinely exists within the evidence?
And how much is created by the categories we introduce before analysis even begins?
That question eventually became the foundation of this paper.
For decades archaeologists have attempted to reconstruct prehistoric landscapes by identifying camps, workshops, quarries, activity areas and settlement systems. These concepts have proved enormously valuable. At the same time, they inevitably shape how we analyse archaeological evidence. Once categories are defined, they often become the framework through which every new observation is interpreted.
While working with regional datasets, I began wondering whether there was another way to approach the problem.
Instead of asking where camps or quarries were located, could we first ask whether the archaeological record itself contained recurring patterns without imposing those labels?
It sounds like a small change.
In practice, it changed the entire research programme.
The technical part of the project appeared, at first, to be the biggest challenge. Selecting environmental variables, preparing GIS layers, writing R scripts and testing different clustering algorithms required months of work. Yet those tasks gradually became routine.
The difficult part turned out not to be computational.
It was conceptual.
Every result forced another question.
What exactly does a cluster represent?
Can an algorithm discover behaviour?
Or does it simply identify statistical similarity?
If prehistoric landscapes are shaped by repeated occupations, erosion, preservation bias and thousands of years of accumulated activity, then similar spatial patterns may emerge for very different reasons.
That meant the algorithm could never provide the interpretation.
Only archaeology could do that.
This became the turning point of the project.
Originally, I thought the clusters would be the discoveries.
Eventually, I realised they were something much more useful.
They were hypotheses.
They were structured questions produced by the data that archaeologists still needed to evaluate critically.
Once I adopted that perspective, the role of machine learning changed completely.
Rather than replacing archaeological reasoning, it became a way of organising uncertainty.
That distinction became one of the central ideas of the paper. Instead of treating computational output as behavioural truth, the framework interprets clusters as inferential propositions that require archaeological explanation. The objective was not to automate interpretation but to make the reasoning behind interpretation more transparent.
Another surprise emerged during validation.
Like many researchers, I hoped that different analytical methods would converge on a single, definitive solution. Instead, the elbow method, silhouette analysis and sensitivity tests occasionally pointed in different directions. Initially, this felt disappointing.
Later, I realised that these disagreements were among the most valuable results.
Science rarely advances because one method provides certainty.
It advances because different forms of evidence begin to converge while also revealing where uncertainty remains.
That changed how I interpreted the entire workflow.
The emphasis shifted away from finding the "best" clustering solution and towards identifying patterns that remained stable across different analytical approaches. Robustness became more important than optimisation.
Working on this study also changed how I think about artificial intelligence in archaeology.
Much of the current discussion focuses on prediction: finding sites, classifying artefacts or automating analysis.
But archaeology is fundamentally a science of incomplete evidence.
We can never observe the behaviours we are trying to explain directly.
Our central challenge is therefore not prediction.
It is inference.
How do we move responsibly from fragmentary observations to credible explanations of the human past?
That question extends well beyond machine learning.
Although this paper focuses on open-air Palaeolithic landscapes in central India, I hope its broader contribution lies in encouraging archaeologists to think more explicitly about the logic of inference itself. Computational methods are becoming increasingly powerful, but their greatest value may not lie in producing more predictions. It may lie in making our reasoning more transparent, our assumptions more visible and our interpretations more accountable.
Looking back, I realise that this paper did not simply produce a new analytical framework.
It changed the questions I now ask as an archaeologist.
Instead of asking "What does this pattern mean?", I now find myself asking something that comes first.
"Why do I believe this pattern represents behaviour at all?"
For me, that is where better archaeological science begins.