A learning automaton selects an action, receives uncertain feedback from its environment, and adjusts the probabilities of its possible actions. The mathematical question is whether a particular update rule improves performance or converges under the assumptions made about that environment. From Tsetlin’s 1961 work to Narendra and Thathachar’s 1974 survey, researchers developed this model as a focused theory of learning in unknown random environments—not as an early version of modern deep reinforcement learning.
What is a learning automaton?
A learning automaton is a decision mechanism connected to an environment that responds uncertainly to its actions. The automaton maintains probabilities over available actions; after an action, feedback from the environment informs an update to those probabilities. Repeated interaction can shift probability toward actions associated with more favorable responses, depending on the update rule and the assumptions of the model. Narendra and Thathachar’s 1974 survey describes stochastic automata operating in unknown random environments as models of learning and focuses on how environmental inputs change action probabilities: their 1974 survey.
| # | Preview | Product | Price | |
|---|---|---|---|---|
| 1 |
|
Introduction to Automata Theory, Languages, and Computation | $21.92 | Buy on Amazon |
| 2 |
|
Automata Theory | $19.91 | Buy on Amazon |
| 3 |
|
Automata Theory: An Algorithmic Approach | $59.28 | Buy on Amazon |
| 4 |
|
Theory of Finite Automata With an Introduction to Formal Languages | $40.81 | Buy on Amazon |
| 5 |
|
The Jazz Theory Book | $49.00 | Buy on Amazon |
It helps to keep the two parts distinct: the automaton selects and updates; the environment supplies uncertain responses. Learning is not simply a claim that the system gets better. It is analyzed through performance criteria and the behavior of action probabilities over time.
How does learning work in a stochastic environment?
- Select an action. The automaton chooses from its available actions according to its current probability distribution.
- Receive environmental feedback. The environment responds, but because it is stochastic, the response to an action is uncertain rather than guaranteed.
- Update action probabilities. A reinforcement or updating scheme uses the feedback to revise the probabilities governing later choices.
- Evaluate behavior over repeated interaction. The analysis asks how performance and action probabilities evolve, and whether the rule meets a specified criterion under the environment’s assumptions.
The update rule is central: different rules can produce different behavior, and conclusions about improvement or convergence are conditional, not automatic. The 1974 survey treats behavior norms, update-scheme design, convergence of action probabilities, and interactions among multiple automata as distinct theoretical questions.
Recommended Free Tools
#1 Best Overall
- Pearson
- Introduction to Automata Theory, Languages, and Computation
How the field developed from 1961 to 1974
1961: Tsetlin’s early work
A 1983 retrospective by Masato Baba attributes the first introduction of learning automata operating in an unknown random environment to Tsetlin in 1961. Baba reports that Tsetlin studied deterministic automata and showed asymptotic optimality under some conditions. This is a later historical attribution; the original 1961 paper was not directly examined here, so the result should not be generalized beyond that account.
1963: stochastic automata
The same retrospective attributes early findings that stochastic automata also have learning properties to Varshavskii and Vorontsova in 1963. As with the 1961 milestone, this chronology rests on the retrospective rather than direct review of the original paper.
Rank #2
- Book - automata theory
- Language: english
- Binding: paperback
1974: a shared theoretical framework
Narendra and Thathachar’s 1974 survey brought theoretical questions and applications into a shared framework. It organized prior work around reinforcement schemes, performance norms, convergence, optimization, hypothesis testing, and interactions among automata. A later overview indexed by PubMed describes the survey as popularizing the label “learning automata” for models introduced in the 1960s; it should be understood as a synthesis and popularization, not the origin of every underlying idea. The later historical overview.
What counts as learning—and what does not follow automatically?
In this field, a learning claim depends on the chosen performance norm, the probability update rule, and the assumptions about environmental responses. A statement that action probabilities converge is not, by itself, enough to establish that a method is optimal; nor does a successful result under one set of assumptions establish the same result for another environment. The 1974 survey’s separate treatment of norms, update schemes, and convergence reflects these distinctions.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Rank #3
The stochastic-environment model is narrower than modern reinforcement learning: it describes an automaton adapting its action probabilities through environmental feedback. The 1961–1974 account does not, by itself, establish a direct lineage to modern deep reinforcement learning or justify applying later terminology to the early work.
How to compare learning-automaton approaches
There is no supported basis here for ranking individual reinforcement algorithms. A useful comparison should instead establish the assumptions and criterion each method uses:
- Feedback model: What response can the environment return, and how is uncertainty represented?
- Update rule: How does feedback change the probabilities assigned to actions?
- Performance criterion: What does the analysis count as successful behavior?
- Convergence or expediency: What property is proved or evaluated, and under which conditions?
- Environment behavior: Is the environment treated as stationary or changing?
These axes help identify whether two approaches address the same problem. Precise comparisons of particular algorithms require their underlying papers and stated conditions.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Further reading
For a book-length follow-up beyond this period, see Narendra and Thathachar’s Learning Automata: An Introduction (1989), cited in a later Wiley chapter on learning automata: the Wiley chapter.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchPC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Quick Recap
Best Value
- The Jazz Theory Book
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

