School of hard knocks: Curriculum analysis for Pommerman with a fixed computational budget

Shelke, Omkar; Meisheri, Hardik; Khadilkar, Harshad

doi:10.1145/3493700.3493709

Computer Science > Artificial Intelligence

arXiv:2102.11762 (cs)

[Submitted on 23 Feb 2021 (v1), last revised 24 Feb 2021 (this version, v2)]

Title:School of hard knocks: Curriculum analysis for Pommerman with a fixed computational budget

Authors:Omkar Shelke, Hardik Meisheri, Harshad Khadilkar

View PDF

Abstract:Pommerman is a hybrid cooperative/adversarial multi-agent environment, with challenging characteristics in terms of partial observability, limited or no communication, sparse and delayed rewards, and restrictive computational time limits. This makes it a challenging environment for reinforcement learning (RL) approaches. In this paper, we focus on developing a curriculum for learning a robust and promising policy in a constrained computational budget of 100,000 games, starting from a fixed base policy (which is itself trained to imitate a noisy expert policy). All RL algorithms starting from the base policy use vanilla proximal-policy optimization (PPO) with the same reward function, and the only difference between their training is the mix and sequence of opponent policies. One expects that beginning training with simpler opponents and then gradually increasing the opponent difficulty will facilitate faster learning, leading to more robust policies compared against a baseline where all available opponent policies are introduced from the start. We test this hypothesis and show that within constrained computational budgets, it is in fact better to "learn in the school of hard knocks", i.e., against all available opponent policies nearly from the start. We also include ablation studies where we study the effect of modifying the base environment properties of ammo and bomb blast strength on the agent performance.

Comments:	8 pages, Submitted to ALA workshop 2021
Subjects:	Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Multiagent Systems (cs.MA)
Cite as:	arXiv:2102.11762 [cs.AI]
	(or arXiv:2102.11762v2 [cs.AI] for this version)
	https://doi.org/10.48550/arXiv.2102.11762
Journal reference:	CODS-COMAD 2022: 5th Joint International Conference on Data Science & Management of Data (9th ACM IKDD CODS and 27th COMAD)
Related DOI:	https://doi.org/10.1145/3493700.3493709

Submission history

From: Hardik Meisheri [view email]
[v1] Tue, 23 Feb 2021 15:43:09 UTC (4,735 KB)
[v2] Wed, 24 Feb 2021 07:54:32 UTC (4,735 KB)

Computer Science > Artificial Intelligence

Title:School of hard knocks: Curriculum analysis for Pommerman with a fixed computational budget

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Artificial Intelligence

Title:School of hard knocks: Curriculum analysis for Pommerman with a fixed computational budget

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators