E420: Sharon Carroll - "A Real Look at Reducing Reinforcement"
FenziFenzi Dog Sports Podcast · 팟캐스트 · 32분
반려견 훈련에서 강화 줄이기는 행동의 신뢰성을 유지하면서 보상 빈도를 전략적으로 조절하는 핵심 과정이다. 훈련 초기 단계에서는 뇌가 신호와 행동, 보상 사이의 연관성을 파악하도록 연속 강화 계획을 유지하며, 행동이 습관으로 자리 잡으면 간헐적 강화 계획으로 점진적으로 전환한다. 보상을 너무 급격히 줄이면 개는 혼란이나 좌절감을 느끼며 행동의 질과 정확도가 저하되므로, 항상 한 번의 반복을 건너뛰는 방식의 테스트를 통해 행동의 견고함을 먼저 확인해야 한다. 또한 개별 행동마다 필요한 강화 비율은 견종의 특성, 동기 부여 수준, 환경적 난이도에 따라 다르므로 각 행동을 별도로 다루어야 한다. 훈련 종료 신호를 명확히 설정하면 반려견은 보상을 받지 못하는 상황에서도 스트레스를 받지 않고 안정적으로 다음 과제로 넘어갈 수 있다. 훈련사들은 최종적인 목표에만 집중해 보상을 성급하게 제거하기보다는, 충분한 강화 이력을 쌓은 뒤 체계적이고 점진적인 단계를 밟아 간헐적 보상 체계로 나아가야 한다.
E420: 샤론 캐럴(Sharon Carroll) - "보상 줄이기에 대한 현실적인 시각"
E420: Sharon Carroll - "A Real Look at Reducing Reinforcement"
0:00
안녕하세요, 멜리사 브로입니다. 여러분은 프렌지 도그 스포츠 팟캐스트를 듣고 계십니다. 이 방송은 프렌지 도그 스포츠 아카데미에서 제공하며, 이곳은 수준 높은 교육을 제공하기 위해 만들어진 온라인 학교입니다. 경쟁적인 도그 스포츠를 위해 가장 최신의 진보적인 훈련 방법만을 사용합니다. 오늘 저희는 샤론 캐롤과 함께 강화 줄이기에 대해 이야기해 볼 것입니다. 안녕하세요 샤론, 팟캐스트에 다시 오신 걸 환영합니다. 안녕하세요 멜리사, 다시 불러주셔서 감사합니다. 다시 오게 되어 정말 기뻐요. 한동안 출연하지 못했거든요. 네, 이번 대화가 정말 기대되네요. 아주 멋진 주제가 될 것 같아요. 이야기 나누게 되어 설렙니다. 네, 시작하기 전에 청취자분들을 위해 간단히 자기소개 좀 해주시겠어요? 물론이죠. 저는 호주 동부 해안 뉴캐슬에 거주하는 전문 반려견 훈련사이자 행동 컨설턴트입니다. 국제 동물 행동 컨설턴트 협회(IAABC)에서 다종 동물 대상 공인 행동 컨설턴트 자격을 보유하고 있습니다. 저는 동물 생물학 및 행동학 분야에서 학술적 자격을 갖추고 있으며, 동물 과학 석사 학위를 가지고 있습니다. 저는 프렌지 교수진 중 한 명이며, 여러 마리의 반려견과 함께 스포츠에 참여하고 있습니다. 랠리 오비디언스, 힐워크 투 뮤직, 뮤지컬 프리스타일, 트릭, 센트 워크 등 다양한 스포츠에 참가합니다. 멋지네요. 좋아요, 그럼 도입부에서 언급했듯이, 오늘은 강화 줄이기에 대해 이야기해보고 싶습니다. 큰 질문 하나 드리죠. 우리가 훈련사로서 강화 줄이기 과정을 이해하는 것이 얼마나 중요할까요? 그리고 이것이 여러분이 강아지에게 가르치는 모든 행동에 대해 시간을 투자해서 해야 하는 일인가요, 아니면 어떤 행동들은 항상 많은 보상을 주기로 계획하는 행동도 있나요? 네, 알겠습니다. 저는 이게 매우 중요하다고 생각합니다. 그 과정을 이해하는 것이 정말 중요하다고 생각해요. 이 팟캐스트를 진행하면서 더 자세히 이야기하게 되겠지만, 항상 시간을 들여 강화 줄이기를 해야 하는지에 대해서는 말이죠. 모든 행동에 대한 강화는 최종 목표에 따라 달라질 것입니다. 그래서 대부분의 스포츠 종목에서는 강화를 줄이는 것이 필요할 것입니다. 왜냐하면 많은 스포츠들이 경기장 안에서 간식이나 장난감을 가지고 있지 않도록 요구하기 때문입니다. 예를 들어 센트 워크(scent work)의 경우, 그런 알림 행동에 대해서는 강화를 줄일 필요가 없기 때문에 우리는 그것을 연속 강화 계획으로 유지할 수 있습니다. 개가 올바른 알림을 수행할 때마다 말이죠. 거의 모든 다른 시나리오에서는 행동을 간헐적 강화 계획으로 전환해야 할 필요성이 있을 것입니다. 즉, 일대일 대응인 연속 강화가 아닌 다른 방식, 다시 말해 여러분이
this is melissa bro and you're listening to the frenzy dog sports podcast brought to you by the frenzy dog sports academy an online school dedicated to providing high quality instruction for competitive dog sports using only the most current and progressive training methods today we'll be talking to sharon carroll about reducing reinforcement hi sharon welcome back to the podcast hi melissa thanks for having me again i'm really glad to be back i haven't been on for a little while and yeah really looking forward to the chat should be an awesome topic i'm excited to talk about it yeah to uh start us out you want to just remind everybody a little bit about kind of who you are sure um i'm a professional dog trainer and behavior consultant based in newcastle on the east coast of australia certified behavior consultant across multiple species with the international association of animal behavior consultants i have some academic credentials in animal biology and behavior including a master's in animal science i'm one of the instructors on the frenzy faculty uh and i compete with some dogs so i compete in lots of different sports uh rally obedience heel to music musical freestyle tricks and scent work fabulous all right so as i kind of mentioned in the intro um i wanted to talk about reducing reinforcement today so big question how important is it for us as trainers to kind of understand the process of reducing reinforcement and then is it something that you take the time to do for every behavior you teach your dog or are there some behaviors you know that you just always plan to reward heavily yeah okay so i think it's hugely important i think it's hugely important that we understand the process um and we're going to i'm sure talk about that a little bit more as we go through this podcast um but in terms of do we always take the time to reduce reinforcement for every behavior that's really going to depend on the end goals so for most sport behaviors it's going to be necessary to reduce reinforcement because a lot of our sports require us to not have treats and toys in the ring with us uh scent work for example though um we don't have to because there's no sort of need to reduce reinforcement for those alerts we can um keep that on a continuous schedule of reinforcement every time the dog performs a correct alert uh in almost every other scenario though there's going to be some need to shift a behavior onto an intermittent schedule of reinforcement meaning something other than that one for one that continuous schedule where you're
2:36
행동을 수행할 때마다 간식을 주는 연속 강화가 아닌 다른 방식이 필요하다는 뜻입니다. 심지어 단순히 신뢰할 수 있는 반려견 행동을 만들기 위한 경우에도 우리는 보통 간헐적 강화 계획으로 전환하기를 원합니다. 왜냐하면 행동이 연속 강화 계획에 있을 때, 보상을 몇 번 연속으로 놓치기만 해도 해당 행동의 질과 신뢰도가 급격히 저하되는 것을 볼 수 있기 때문입니다. 하지만 일단 행동을 간헐적 강화 계획으로 옮겨놓으면 우리는 선택권을 갖게 됩니다. 여전히 매우 높은 강화 빈도로 보상할 수 있지만, 만약 몇 번의 행동 반복에 대해 강화를 놓치더라도 행동의 질이나 신뢰도가 저하되는 결과로 이어지지는 않을 것입니다. 물론 행동들 중에는 분명히 매우 높은 강화 빈도로 유지되어야 하는 행동들이 있을 것입니다. 우리가 행동의 많은 수행을 강화해 주는 경우죠. 반면에 강화 빈도가 매우 낮아도 되는 행동들도 있는데, 즉 수행의 아주 작은 비율만 강화되거나, 심지어 강화가 제공되는 경우도 있을 수 있습니다. 우리 외의 다른 방식으로부터 보상이 주어질 수 있기 때문에 행동에 대한 보상이 전혀 필요 없는 것처럼 보일 수 있습니다. 간식이나 장난감을 계속 줄 필요가 없기 때문이지만, 실제로는 다른 방식으로 그 행동이 유지되고 있는 것입니다. 음, 다른 방식으로 신뢰성이 유지되고 있다는 점이 흥미롭네요. 정말 어떻게 보면 생각해 볼 만한 흥미로운 방식이네요. 나중에 다시 언급하겠지만, 어떻게 결정하시나요? 행동이 특정 수준에 도달했을 때 어떤 행동에 대한 보상을 줄일지 결정하는 기준이 무엇인가요? 어느 시점에 이 행동이 충분히 확실하다고 판단해서 보상을 줄이기 시작하시나요? 아니면 행동이 확실하게 자리 잡기 전부터 시작하시나요? 잘 모르겠네요. 그 질문에 대해서는 과학적인 답변이 있습니다. 뇌에서 실제로 무슨 일이 일어나는지와 관련된 것이죠. 그리고 눈앞에서 보이는 행동을 직접 평가하는 좀 더 실용적인 답변이 있습니다. 반려견이 새로운 기술이나 행동을 배울 때 실제로 세 가지 단계가 나타납니다. 첫 번째는 습득 단계인데, 이때는 뇌가 무슨 일이 일어나고 있는지조차 제대로 파악하지 못할 때입니다. 개는 무엇이 보상으로 이어질지 전혀 모르는 상태입니다. 그저 추측할 뿐이고, 무작위로 행동을 해보며 무엇이 보상을 받고 무엇이 받지 못하는지 알아내려고 노력하는 중이죠. 그래서 개는 전혀 감을 잡지 못하고 뇌는 필사적으로 어떤 연관성을 찾으려고 합니다.
giving a treat every time they perform the behavior um even if it's just to produce a reliable pet dog behavior we still really usually want to uh shift to that intermittent schedule because when a behavior is on a continuous schedule then when we miss a reward just even for several repeats of the behavior in succession uh we can see this rapid deterioration in the quality and reliability of that behavior but once we've put that behavior onto an intermittent schedule of reinforcement then we have options we can still reward on the really high rate of reinforcement but if we do miss reinforcing for several reps of uh that behavior in succession it won't result in any deterioration in the quality or reliability of the behavior um there absolutely will be behaviors though that need to be maintained on that very high rate of reinforcement um where we sort of reinforcing many of the performances um of the behavior and then other behaviors where the rate of reinforcement can actually be very low so only a very small percentage of the performances are reinforced uh or potentially even the reinforcement will be provided in a way other than from us so it may appear like the behavior doesn't need reinforcing at all because we don't need to keep handing over treats and toys but really it's just that it's being maintained through some other way um the reliability is being maintained some other way interesting it's really kind of an interesting way to think about it um i'm sure we'll circle back to that but how do you kind of decide for those behaviors that you are going to reduce reinforcement when the behavior is at that point like when do you look at something and go okay this is a solid enough behavior that it's time to start reducing reinforcement or maybe you start even before it's a solid behavior i don't know yeah well there's a scientific answer to that question uh where we're referring to what's actually happening in the brain and then there's the more practical answer where we're just assessing directly from the behaviors we're seeing in front of us so there's actually three phases that occur when our dog's learning a new trick or learning a new behavior there's the acquisition phase and that's when the brain doesn't really even have a clue what's happening the dog doesn't know what's what's going to end up in reinforcement they're just sort of guessing they're just throwing out random guesses and trying to work out what gets rewarded and what doesn't so they have no clue and the brain is madly trying to
4:58
아, 내가 그 행동을 했을 때 보상을 받았구나 하는 연관성을 찾으려 하고 시간이 지나면서 다음 단계인 행동-결과 단계에 도달하게 됩니다. 이때 연관성이 확인되기 시작해서 뇌는 '내가 그 행동을 할 때마다 보상을 받는 게 확실해'라고 판단하게 됩니다. 하지만 그 시점부터 뇌는 그 행동이 수행될 때마다 지속적으로 확인 작업을 거치게 됩니다. 개들은 '내가 보상을 받고 있는 건가?'라고 확인하려 할 텐데, 바로 그것이 뇌에게 그 선택이 옳은 선택이었고, 그 신호에 대한 옳은 행동이었다고 알려주는 것이기 때문입니다. 그것이 옳은 행동이고 보상을 받을 행동이라는 점이 확실해지면 습관 형성 단계로 넘어가게 되는데, 이때 뇌는 비로소 좋아, 이제 연관성이 아주 명확해졌어. 신호가 나타났을 때 보상을 받기 위해 내가 이 행동을 해야 한다는 것을 확실히 이해했어라고 판단하며 매우 명확한 연관성이 형성됩니다. 이제 뇌는 덜 주의 깊게 감시해도 됩니다. 즉, '그래, 이제 내가 맞다는 걸 알아'라는 상태가 되어서 이 행동을 수행할 때마다 매번 감시할 필요가 없는 것이죠. 그래서 행동은 개들이 의식적으로 어떤 행동을 해야 할지 결정하고 깊이 생각할 필요 없이, 오직 신호에 대한 반응으로만 수행될 것입니다. 뇌는 그것이 옳은 선택임을 꽤 확신하고 있으므로 더 이상 계속 감시하지는 않을 것입니다. 아니, 그렇게 밀착 감시하지는 않더라도 여전히 감시는 유지하지만, 매번 반복할 때마다 하는 것은 아니죠. 실질적인 관점에서 보면 초기 단계에서는 많은 추측이 보이고, 맞는 경우도 있고 틀리는 경우도 있습니다. 우리는 여전히 유도하거나 셰이핑(shaping)을 하고 있고 도구를 사용하고 있는데, 지금은 결코 보상을 줄여서는 안 되는 시기입니다. 다음 단계로 넘어가면 정확한 반응이 많이 나타나는 것을 볼 수 있을 겁니다. 하지만 보상을 건너뛰면 다음 반복에서 개의 행동에 즉각적인 변화가 있을 수 있습니다. 반면 마지막 단계에 도달하면 실제로는 행동이 신호에 반응하여 신속하고 정확하게 나타나는 것을 볼 수 있습니다. 즉, 행동이 수행되기 전에 보상이 보이지 않더라도 행동을 수행한 후에는 매번 보상을 주면서 중요한 점은 그 시점에 다다랐을 때 비로소 보상을 건너뛰는 것을 고려하기 시작할 수 있다는 것입니다. 반복 훈련 중 하나를 마친 후 다음 반복 훈련에 영향이 없다면 우리는 강화 줄이기 과정을 시작할 수 있습니다. 저는 강화 줄이기를 이런 식으로 생각하는 걸 좋아합니다. 실제로 테스트를 해볼 수 있다는 거죠. 과연 강화를 줄였을 때 행동이 변하는지 확인하고, 만약 변하지 않는다면 행동이 견고하게 자리 잡았다고 볼 수 있습니다. 우리는 과정 중
identify some sort of association between oh when i did that behavior i got a reward and then over time we're going to reach that next phase where we get that action outcome phase where the associations are now being identified so the brain's going i'm pretty sure every time i do that behavior i get that reward but the brain's going to keep constantly monitoring at that point every time that behavior is performed they're going let me check am i getting that reward because that's what's telling the brain that that was the correct choice that was the correct behavior for that cue that is the correct behavior and that's what's going to get rewarded then we move through to this habit formation phase and that's where the brain actually goes okay the association is really clear now i do understand that when that cue happens i should do this behavior in order to get the reward and so this is really clear associations being formed and the brain's now able to monitor less closely so it's sort of like okay i know that i'm right now i don't need to monitor every time this behavior is performed um so it's more um the behavior is going to be performed just in response to the cue without as having without the dog having to consciously really make a decision and really think about what behavior to perform and um the brain's sort of pretty sure that that's the right choice so it's it's not going to keep monitoring now from a practice or not monitor as closely it still keeps monitoring but just not every rep and so from a practical perspective in the early phase we just see lots of guessing some correct some incorrect uh we're still doing the luring or the shaping we're using props that's certainly not a time to reduce reinforcement in that next phase we're going to see lots of accurate responses happening um but when we skip a reward there may still be an immediate change to our dog's behavior in that next rep whereas when we reach that end phase what actually happens now is the behavior is occurring rapidly and accurately in response to our cue we see that happening so without rewards being apparent before the behavior is performed but still giving a reward after every performance of the behavior but importantly when we can sit at that point we can start considering skipping a reward after one of the reps and if the next rep is unaffected then we can start the process of reducing reinforcement i like that way of thinking about it of actually being able to test okay does the behavior change as we reduce reinforcement and if not then we've kind of gotten it solid we're at
7:30
올바른 단계에 있는 거죠. 정확히 말하면, 항상 한 번의 반복 훈련을 건너뛰는 것을 기억해야 합니다. 한 번 반복 훈련을 건너뛰고 다시 보상을 주고, 또 한 번 건너뛰고 다시 돌아가는 식이죠. 만약 여러 번의 반복 훈련을 건너뛰기 시작하면 행동이 무너지는 현상을 볼 수도 있습니다. 어쨌든 딱 그 한 번의 반복 훈련이 우리에게 정보를 줍니다. 우리는 그 한 번을 건너뛰고 다음에 큐를 줬을 때 어떤 일이 일어나는지 지켜보는 것입니다. 강화 줄이기에 어떻게 접근하는지 전체적인 개요를 말씀해 주실 수 있나요? 테스트를 완료하고 한 번의 반복 훈련을 했는데도 행동이 저하되지 않았다면, 그 다음은 무엇인가요? 네, 맞습니다. 그 시점부터 시작해서 행동이 확립되었는지 확인합니다. 여기서 우리가 확인해야 할 몇 가지는 행동이 수행되기 전에 간식이나 장난감이 필요하지 않아야 한다는 점입니다. 이건 사실 강화 줄이기와는 별개의 문제일 수 있지만, 우리가 행동의 신뢰성을 얻으려고 할 때 발생하는 문제들을 보면 대부분 아직 그 단계에 도달하지 못한 경우가 많기 때문입니다. 그래서 행동이 아직 확실하게 이루어지지 않은 상태에서는 절대 강화 줄이기를 시작하지 않습니다. 우리가 행동에 큐를 주기 전에 간식이나 장난감 없이도 큐에 반응하여 정확하고 신뢰성 있게 행동이 나타나야 합니다. 그것이 첫 번째 단계입니다. 일단 그 행동이 빠르게 그리고 정확하게 나타난다는 것을 알게 되면, 큐를 주기 전에 간식이나 장난감이 필요 없이 첫 번째 큐에 바로 반응하므로, 이제 그 한 번의 보상을 건너뛸 수 있습니다. 이제 무슨 일이 일어나는지 분석합니다. 좋습니다. 행동이 그대로 견고하고 아무런 변화가 없다면 다음 반복 훈련에서도 마찬가지일 것입니다. 이제 모든 것이 동일해졌으니 간헐적 강화 스케줄로 전환하기 시작할 수 있지만 강화 빈도는 여전히 매우 높게 유지합니다. 그래서 당분간은 10번 반복할 때 1번 정도만 보상을 주지 않는 식으로 할 수도 있고, 그다음에는 두 번 연속 보상을 건너뛸 수도 있습니다. 하지만 너무 연달아 두 번을 건너뛰지는 말고, 보상을 주고 안 주는 식으로 번갈아 가며 진행할 수 있겠죠. 그런 식으로 점진적으로 단계를 낮춰가며, 결국에는 스펙트럼의 반대편 끝까지 가서 보상을 주지 않고도 9번이나 10번의 행동을 요구할 수 있는 수준에 도달할 수 있습니다. 하지만 그 과정은 반드시 점진적으로 이루어져야 합니다. 그 과정 중에 만약 훈련을 한 번 실패하거나, 큐에 대한 반응이 지연되거나 정확도나 열의에 변화가 보인다면, 일단은 하던 일을 잠시 중단하고, 매번 보상을 주는 단계로 되돌아가는 것이 좋습니다. 흥미롭네요. 좋습니다. 그렇다면 훈련에서 간식이나 장난감을 제거하는 과정을 진행할 때, 너무 빠르거나 너무 느리게 진행하고 있다는 것을 어떻게 판단하시나요? 방금 답변해주시긴 했지만,
the right phase in the process exactly and that's remembering that it's always that one rep we skip one rep and we go back to rewarding and we skip one rep and go back to if we start skipping multiple reps we might see some disintegration anyway but just that one rep that's what gives us information we skip that one rep and we look at what happens that very next time we deliver the cue can you give us kind of an overview of how you actually approach reducing reinforcement so you've done your test you got that one rep and the behavior did not degrade what next okay so yeah we start at that point we check the behavior is established so we know the couple of things we want to check here is that we don't need treats or toys present before the behavior is performed now that's separate really to anything to do with reducing reinforcement but a lot of the time when we look at just problems that are occurring with people trying to get reliability from behaviors it's because there's those they haven't even reached that point yet so we don't we definitely don't start reducing reinforcement until the behavior happens reliably and accurately in response to the cue without us having treats and toys present before we cue that behavior so that's the first thing and so once we know that behavior is rapidly and accurately occurring on our first cue without the need for anything any treats or toys before it then we can skip that one reward now we analyze what happens okay great if the behavior just is solid nothing changes the very next rep everything's the same now we can start to shift to that intermittent schedule but we still keep that rate of reinforcement very high so we might be only missing one out of every 10 reps for a little while and then we can start skipping two but maybe not even two in a row too close together maybe do one with with a reward one without one with a reward one without and then we might gradually move down to where we're on the opposite sort of end of the spectrum where we can ask for nine or ten behaviors before we even need to give a reward but we do need to do that process gradually in that in that process though if we miss a rep and we see any delay in response to the cue or any change to the accuracy or enthusiasm we just sort of abort the mission temporarily and go back to rewarding every time interesting all right so how do you decide i know you kind of answered this there but you know if you're moving too quickly you're not quickly enough when you're moving you know the food
9:55
너무 빨리 혹은 너무 느리게 진행하지 않도록 어떻게 확인하거나 모니터링하시나요? 네, 너무 빨리 진행하고 있다면, 개에게서 혼란스러워하거나 좌절하는 징후가 보일 것입니다. 그 좌절감은, 개별 강아지에 따라 다르겠지만, 훈련을 포기하는 결과로 나타날 수 있습니다. 그냥 가버리거나, 이제는 더 이상 하고 싶지 않다는 태도를 보이거나, 아니면 흥분해서, 우리에게 짖거나 옷을 물어뜯거나 방 안을 미친 듯이 뛰어다니는 행동(줌이)을 할 수도 있습니다. 따라서 혼란이나 좌절의 징후가 보인다면, 강화 빈도를 너무 빨리 줄인 것입니다. 학습 과정이 끝나기 전에 시작했거나, 보상 빈도 자체를 너무 급격하게 낮추려고 한 것이죠. 실제 보상 횟수를 너무 빨리 줄인 경우이며, 이럴 때는 신뢰도도 떨어지는 것을 보게 될 것입니다. 가끔은 행동을 유도하기 위해 반복적인 신호가 필요할 수도 있고, 때로는 신호에 정확하게 반응하기도 합니다. 어떨 때는 반응하지 않거나, 신호와 행동 사이에 지연 시간이 발생할 수도 있어서 우리가 신호를 주면 그 행동을 하기까지 잠시 멈춤이 발생합니다. 이 모든 것들은 우리가 행동이 습관 형성 단계에 도달하여 확립되기 전에 보상을 줄이려고 했거나, 간헐적 강화 계획으로 넘어갔는데 너무 서둘러서 보상 비율을 충분히 점진적으로 줄이지 않고 너무 급격하게 낮추려 했다는 것을 알려줍니다. 다행히도 보상 줄이기를 시작하기에 너무 늦은 시점이라는 것은 없습니다. 많은 사람이 너무 빨리 시작할 수도 있다고 생각하지만, 어쩌면 너무 늦게 시작할 수도 있지 않을까 생각할 수 있습니다. 즉, 연속적 강화 계획을 너무 오래 유지해서 개가 그것에 의존하게 된 것은 아닐까 하고 말이죠. 하지만 실제로는 그런 일이 일어나지 않습니다. 너무 오래 기다린 것이 아니라, 사람들이 갑작스럽게 변화를 주기 때문입니다. 따라서 얼마나 오래 기다렸든 간에, 행동을 오랫동안 연속적 강화 계획으로 유지해 왔다면, 여전히 점진적으로 간헐적 강화 계획으로 전환해야 하며, 갑작스럽게 전환하려고 해서는 안 됩니다. 또는 보상 비율을 너무 빨리 낮추려고 하지 마십시오. 문제가 발생하는 지점이 바로 그때이기 때문입니다. 음, 여전히 체계적이고 점진적으로 진행해야 합니다. 이 시점에서 제가 계속해서 연속적 강화 계획과 간헐적 강화 계획이라고 말하고 있는데, 이 개념을 짚고 넘어가는 것이 좋을 것 같습니다. 연속적 강화 계획이란 일대일 방식을 의미하며, 개가 행동을 할 때마다 간식을 받는 것입니다. 간헐적 강화 계획은 행동을 할 때마다 보상을 주는 것이 아니라는 뜻입니다.
or the toys from the equation how do you kind of keep an eye on that or like how do you monitor that you're not moving too fast or not moving quickly enough yeah well if you're moving too quickly then we'll see signs of confusion and frustration we'll see and that frustration depending on the individual dog that frustration might result in quitting it might result in them just walking away going i don't want to do this anymore or it may result in them getting amped up they might be barking at us mouthing our clothing doing zoomies around the room so if we're seeing signs of confusion or frustration we've tried to reduce reinforcement too quickly either we've we've started before the learning process was complete or we're just trying to get that rate of reinforcement down too quickly like the actual rate of rewarding down too quickly and also we're going to see reduced reliability so maybe they need repeated cues to get the behavior sometimes they're responding accurately to the cues sometimes they're not or there might even be a lag time between the cue and the behavior so we give the cue there's a pause before they do the behavior all of those things would tell us that either we've started to try to reduce reinforcement before that behavior is established before it's reached habit formation phase or we've got onto that intermittent schedule but we've been in too much a rush to get down to a much lower rate of reinforcement instead of sort of doing that quite gradually now fortunately though there's never a point where we've waited too long before starting reducing reinforcement so a lot of people think it's like oh you know you could beat um you can start it too quickly but maybe you could also start it too late like you've stayed on that continuous schedule of reinforcement for so long and now your dog's reliant on it but that's not actually what happens it's not that we've waited too long it's that people then sudden do a sudden shift so no matter how long we've waited if we've kept that behavior on a continuous schedule for a long while then we need to incrementally still shift onto that intermittent schedule not try to make that shift really suddenly or try to drop that rate of reinforcement really quickly because that's when the problems actually occur um we still need it to be systematic and incremental um and i think at this point it's probably worth talking about like i keep saying continuous schedule an intermittent schedule continuous schedule of reinforcement just means one for one every time they do the behavior they get a treat
12:23
즉, 때로는 행동을 하고 간식을 받기도 하지만, 때로는 행동을 해도 간식을 받지 못할 때도 있다는 의미입니다. 이번에는 간식을 주는데 인간의 관점에서 보면 음, 흔히 이렇게 이야기하곤 합니다. 연속 강화 계획은 자판기와 같다고요. 우리가 다가가서 돈을 넣고 버튼을 누르는 행동을 하면, 매번 간식이 나올 것을 기대하게 되죠. 왜냐하면 그게 바로 연속 계획이니까요. 전등 스위치도 마찬가지입니다. 방에 들어가서 불을 켜면, 불이 켜지는 강화가 따를 것이라 예상합니다. 스위치를 켜는 행동에 대해 우리가 원하는 결과이기 때문이죠. 연속 강화 계획은 새로운 행동을 학습할 때 아주 좋습니다. 왜냐하면 뇌가 큐(신호), 행동, 그리고 보상 사이의 연관성을 파악하기가 매우 쉽기 때문입니다. 보상이 매번 제공되기 때문이죠. 하지만 보상이 예상되기 때문에 소거 또한 매우 빠르게 일어납니다. 그래서 만약 연속해서 너무 많이 보상을 건너뛰게 되면 그 행동은 소거될 것입니다. 행동이 악화되어 양질의 행동을 기대하기 어렵게 되고 신뢰도가 떨어지게 됩니다. 음, 또한 연속 강화 계획에서는 보상이 제공되지 않을 때 좌절로 인한 행동이 나타날 위험이 매우 높습니다. 사람들이 자판기를 두드리는 것을 본 적이 있을 겁니다. 왜냐하면 그들은 간식을 원하니까요. 간식은 보장되었고 당연히 나올 것이라 예상했는데 나오지 않았기 때문이죠. 그래서 이제 우리는 이런 반응을 얻게 됩니다. 개들도 마찬가지입니다. 만약 우리가 연속 강화 계획을 유지하면서 보상을 주지 않거나 처음 한두 번 이상 계속 건너뛰게 되면, 그들에게는 그 상황이 힘들 수 있습니다. 그래서 우리는 좌절로 인한 행동을 보게 될 수도 있습니다. 자, 이제 간헐적 강화 계획에 대해 이야기해 보죠. 많은 사람들이 인간의 관점에서 이를 슬롯머신에 비유합니다. 돈을 넣고 버튼을 누르지만, 우리는 행동을 할 때마다 항상 보상을 받는 것은 아니라는 것을 알고 있습니다. 자, 간헐적 강화 계획은 음, 지속성을 만들어내고 소거에 대한 저항을 기르는 데 정말 좋습니다. 연속해서 버튼을 여러 번 계속 누르게 될 것입니다. 매번 보상을 받지 못한다고 해서 버튼을 계속 누르고 싶은 욕구가 줄어들지 않기 때문이죠. 이것이 바로 우리가 반려견에게 원하는 것입니다. 즉, 반려견이 큐(신호)를 들을 때마다 그 행동을 수행하는 것이죠. 정말 끈기 있고 확실하게 행동하게 만드는 것입니다. 반려견의 입장에서는 언젠가는 보상이 주어질 것이라고 확신하기 때문입니다. 그것이 바로 우리가 원하는 바입니다. 우리가 연속 강화 스케줄에서 간헐적 강화 스케줄로 전환할 때 일어나길 바라는 일이죠. 우리는 반려견이 언젠가는 보상을 받을 것이라는 이해를 높이도록 돕고 싶을 뿐입니다.
an intermittent schedule just meaning not every time they do the behavior so sometimes they do the behavior and they get a treat like the treat comes sometimes they do the behavior and they they don't get a treat that time and if we look at it from a human perspective um it's often talked of as that continuous schedule of reinforcement being like a vending machine like where we go up we perform the behavior of putting the money in and hitting the button and we expect a treat every time because that's what it's a continuous schedule same with a light switch we walk in the room we flick the light on we expect to see the reinforcement of light because that's what we want for our putting a light switch on continuous schedules are great for learning new behaviors because it's very easy for the brain to identify that association between the cue the behavior and the reward because the reward's happening every time but also extinction occurs really rapidly because the reward is expected so if we if we sort of miss too many in a row we will get extinction of that behavior the behavior will deteriorate and won't get such good quality behavior will lose some reliability um also the risk of frustration-based behaviors when rewards aren't delivered is really high when we're on a continuous schedule so you know how many times people start sort of banging on a vending machine because they want the treat because it was guaranteed it was expected and it should have come and so now we get these and it's the same with our dog if we if we're on that continuous schedule and we are not delivering the rewards and we're skipping more than just one that first time that can be hard um on them so we might see frustration-based behaviors now the intermittent schedule a lot of people in the human terms will refer to that like a slot machine we put the money and we hit the button and we know that sometimes we're going to get rewarded not every time we perform that behavior now intermittent schedules of reinforcement are really great for um producing some persistence some resistance to extinction so you will keep hitting that button multiple times in a row and just because it doesn't get rewarded every time it doesn't deteriorate your desire to keep pushing the button and that's what we want from our dog is we really want that behavior to just every time they hear that cue they perform that behavior really persistently really reliably because in their mind they're still sure at some point it is going to be reinforced and that's what we're trying to get um what we're trying to have happen when we switch from
14:47
여전히 보상은 주어지지만 매번 주어지는 것은 아니라는 점을 말이죠. 그러니 계속 시도하고 행동을 수행하면 언젠가는 보상을 받을 수 있다는 것을요. 저는 인간의 관점에서 이런 사례들을 보면서 개들에게 이것이 얼마나 복잡한 문제인지 깨닫지 못하기 쉽다고 생각합니다. 왜냐하면 인간은 말이나 글로 미리 이러한 상황들을 설명할 수 있기 때문입니다. 누군가가 자판기가 어떻게 작동하는지, 슬롯머신이 어떻게 작동하는지 미리 알려주었을 테니까요. 하지만 동물들은 경험을 통해서만 학습해야 합니다. 따라서 어떤 행동이 연속 강화 스케줄에 있다면 보상에 대한 기대치가 생기게 됩니다. 우리는 그 행동을 간헐적 강화 스케줄로 매우 점진적으로 전환해야 합니다. 그들이 '아, 그렇구나. 한 다섯 번이나 여섯 번에 한 번꼴로 보상을 받는 걸 수도 있겠네'라고 생각할 때까지 말이죠. 어쩌면 10번이나 12번에 한 번일 수도 있고요. 하지만 갑자기 40번으로 늘려버리면 행동의 소거가 발생할 수 있습니다. 우리는 그것을 점진적으로 구축해야 합니다. 습관 형성이라는 것은 행동이 의식적인 판단이나 결정 없이도 큐에 반응하여 나타나는 것을 의미한다는 점을 기억하세요. 뇌는 모니터링을 덜 자주 하게 되지만 여전히 모니터링은 하고 있습니다. 예를 들어, 우리가 방에 들어가서 전등 스위치를 켰는데 불이 켜지지 않는다면 어떻게 될까요? 우리는 계속 스위치를 누르게 될 것입니다. 매번 보상이 주어지지 않는다고 해서 버튼을 누르고 싶은 마음이 사라지지 않기 때문입니다. 그것이 우리가 반려견에게 바라는 점입니다. 네, 전구가 나갔다는 걸 알지만, 그 방에 다시 들어가면 우리는 또 습관적으로 전등 스위치를 켤 겁니다. 습관이 형성되었기 때문이죠. 그래서 우리는 방에 들어가서 전등 스위치를 켜기 전까지 꽤 오랫동안 그렇게 행동합니다. 그러다가 '아 맞다, 불이 안 들어오지'라고 생각하게 되는 거죠. 그러니 소거 현상은 여전히 일어난다는 점을 기억하는 게 중요합니다. 단지 습관 형성 단계에 도달하면 훨씬 천천히 일어날 뿐이죠. 예를 들어, 만약 누군가 당신에게 불이 켜진다는 보장이 없는 버튼을 줬다고 칩시다. 손전등을 줬는데 처음 눌렀을 때는 불이 들어오지 않아요. 두 번째 누르니 불이 들어오고, 그다음 사용할 때는 세 번 눌러야 켜지고, 어떨 땐 두 번, 어떨 땐 다섯 번을 눌러야 하죠. 하지만 그게 작동 방식이라는 걸 알게 되면 당신은 그 버튼을 여러 번 누르기 시작할 겁니다. 그러면 고장 난 전등 스위치나 자판기만큼 좌절하지는 않겠죠. 이런 설명이 연속 강화와 간헐적 강화 사이의 개념을 이해하는 데 도움이 될지 모르겠습니다. 하지만 누군가에게는 이해하는 데 도움이 되었기를 바랍니다. 네, 정말 좋은 설명이었어요. 강화 보상을 줄이려는 구체적인 행동에 따라 이 과정이 달라지기도 하나요?
that continuous schedule to the intermittent is just to help them build up that understanding that it will be still reinforced but it's just not going to be reinforced every time so just persist just keep going just perform the behavior and at some point that behavior will be reinforced and i think it's easy for us from a human perspective to look at those examples and not realize how complicated that is for our dogs because for humans we use language like verbal and written language to explain these things ahead of time someone told you that's how a vending machine works someone told you that's how a slot machine works but animals need to learn that through experience so if a behavior is on a continuous schedule there's an expectation of a reward we need to shift that behavior onto the intermittent schedule very gradually until they go oh okay so maybe just every five or six times it's going to be rewarded maybe every 10 or 12 and we can't suddenly go to 40 still because we're going to get a deterioration we have to build that gradually um so remember just remembering that habit formation really just means the behavior will happen in response to the cue without needing conscious assessment or decision making and the brain monitors less frequently but it still monitors so if we walk into the room into a room and we turn the light switch on and the light doesn't come on okay we work out that it's blown globe but when we go in that room again we're going to turn that light switch on again because it's a habit it's reached habit formation and we do that quite a while before we walk in the room and think before we turn the light switch on we get oh that's right no point the light's not working so it's important to remember extinction still happens it just happens more slowly once we've reached that habit formation phase um and whereas if for example someone gave you a button that wasn't guaranteed to have a light come on so it gives you a torch and you press it the first time torch doesn't work and then you press it a second time torch works and then next time you use it it's like you have to press it three times and then it's two times and then five times but when you know that that's how it works you will start just pushing that button multiple times um to get and you won't be as frustrated as you would be with a light switch or a vending machine that's not working anymore so i don't know if that helps people to understand that that the concept at all between continuous and intermittent but hopefully it helps someone to understand it yeah no i think that i
17:18
근본적으로는 아니에요. 과정은 동일합니다. 다만 어떤 행동은 최종 단계에 도달하기까지 더 오랜 훈련이 필요할 수 있습니다. 어떤 행동은 정말 간단해서 몇 번의 세션 만에 바로 습관 형성 단계에 도달하기도 하죠. 개는 '아, 큐(cue)가 주어지면 행동을 하고 보상을 받는구나'라는 걸 금방 깨닫고, 이제 우리는 강화를 줄이기 시작할 준비가 된 겁니다. 반면에 다른 행동이나 개들의 유형에 따라서는 몇 주, 심지어 몇 달 동안 훈련해야 할 수도 있습니다. 네, 그렇겠네요. 그렇죠. 이제 그 행동에 대한 강화(reinforcement)를 줄이기 시작할 준비가 되었습니다. 그리고 행동마다 유지하기 위해 필요한 강화 비율도 다를 것입니다. 아주 쉬운 행동은 동일한 강화 비율이 필요하지 않을 것이며, 매우 복잡한 행동과도 다를 것입니다. 또한 이는 다음과 같은 요소들에 의해 영향을 받습니다. 개 개별의 타고난 특성 말이죠. 매우 재미있는 일을 하는 에너지가 넘치는 개에게는 매우 다른 강화 비율이 필요할 것입니다. 아주 낮은 에너지의 개에게 높은 노력과 많은 신체적 노력을 요구한다면, 그 행동을 유지하기 위해 상당히 자주 강화하기를 원할 것입니다. 그리고 알다시피 그 활동 자체에 본질적인 가치가 있어서 개가 그 일을 정말 좋아하는 경우라면 행동을 학습한 후에는 굳이 간식이나 장난감을 줄 필요가 없을 것입니다. 아니면 아주 가끔씩만 주면 될 것입니다. 마찬가지로 어떤 것을 아주 좋아하는 개도 훈련의 도전과 복잡성을 좋아하는 개와 그렇지 않은 개가 있기에 우리의 강화 비율은 그러한 것들에 따라 달라질 것입니다. 또한 행동의 익숙함 정도도 중요합니다. 아주 새로운 행동은 여전히 매우 높은 강화 비율로 진행될 수 있지만, 매우 익숙하고 확립된 행동은 매우 낮은 강화 비율로 크게 낮아질 수 있습니다. 심지어 환경의 영향도 있을 수 있겠죠. 우리는 특정 행동의 신뢰도를 높이기 위해 강화 비율을 높이기도 합니다. 왜냐하면 당시 그 개에게는 환경이 더 힘들기 때문이죠. 우리가 많은 경우 강화를 줄이려고 하는 관점에서 생각해보면 꽤 흥미로운 일입니다. 더 어려운 상황에서도 그 행동들이 유지될 수 있도록 하기 위해서 말이죠. 네, 그렇지만 그 부분을 나누어서 생각해야 합니다. 이렇게요. 좋아요, 처음에는 그런 어려운 환경에서 강화 비율을 높였다가 나중에는 결국 물론 낮은 강화 비율로 다시 줄이기는 하지만, 사람들이 더 넓은 환경으로 나갈 때
think that was a good explanation does any of this change depending on what behavior specifically we're looking to reduce reinforcement for uh fundamentally no the process is is the same but some behaviors may need to be trained for longer before we reach that end point so some behaviors are really simple within a couple of sessions that's they're already at habit formation the dog knows oh i do you do the cue i do the behavior i get the reward okay we're ready to start reducing reinforcement but whereas other behaviors or different types of dogs we might be working for weeks and weeks sometimes months before we're really ready to start reducing reinforcement for that behavior um so and and different behaviors will require different rates of reinforcement to maintain them as well so a very easy behavior won't need the same rate of reinforcement as very complex behavior and it's also going to be affected by things such as the innate traits of the individual dog so a high energy dog doing a really fun thing is going to need a very different rate of reinforcement to a very low energy dog being asked to do a high effort high physical effort they're going to want that reinforcement fairly frequently to maintain it and you know same as if there's an intrinsic value in the activity that dog really just loves doing that thing you're not going to need to probably give them treats or toys once they've learned the behavior or you may certainly only need to do that very infrequently um same with a dog that might love the challenge and complexity of work versus a dog that doesn't love that so much so our rates of reinforcement will vary with things like that um also the familiarity of the behavior a very new behavior might still be on a very high rate of reinforcement whereas a very familiar established behavior might be way down on a very low rate of reinforcement um and even the influence of the environment maybe um we increase the rate of reinforcement for a specific behavior just to help increase the reliability of that behavior because the environment's more challenging for that dog at that time it's kind of an interesting thing to think about from the standpoint of a lot of the times we are seeking to reduce reinforcements that we can have those behaviors hold up in the more challenging situations yeah it's like yeah but but we have to we have to separate that in pieces it's like okay initially in that challenging environment we increase the rate of reinforcement then eventually we do decrease back to our lower rate of reinforcement but yeah it is common people go out into those bigger
19:43
간식과 장난감 사용까지 동시에 중단하고 싶어 하는 것은 흔히 있는 일입니다. 맞아요, 그리고 왜 그것이 자신의 목표를 달성하기 어렵게 만드는지 알 수 있을 겁니다. 제가 다음으로 질문하려던 부분인데, 왜 많은 팀이 강화를 제거하는 데 어려움을 겪는다고 생각하시나요? 아마 그게 답변의 일부인 것 같네요. 네, 제 생각에는 때때로 최종 목표에 대해 너무 많이 생각하는 것 같아요. 행동을 가르치기도 전에 이미 이런 걱정을 하는 거죠. 음, '강화가 이렇게 많이 필요하면 안 되는데'라고 생각하지만, 사실은 아니에요. 강아지는 그 시점에 행동을 배우고 있으므로 많은 강화가 필요합니다. 그래서 과정 중에 너무 일찍 보상을 제한하려고 하거나, 행동을 수행할 동기를 충분히 만들지 않는 것이죠. 음, 충분히 긴 강화 이력을 쌓기도 전에 간식과 장난감을 너무 빨리 줄이려고 하는 것 말이에요. 음, 네. 어쩌면 행동을 지시하기 전에 보상을 미리 보여주고 있어서 상황을 방해하는 경우도 많죠. 그래서 강화를 줄이기 전에 먼저 그런 일을 멈춰야 합니다. 따라서 강화 비율을 너무 빨리 줄이는 것이 사람들이 갑자기 문제를 겪는 큰 이유 중 하나이고, 연속해서 너무 많은 보상을 건너뛰는 것도 문제죠. 말씀하신 대로 대회장에서는 흔히 일어나는 일인데, 훈련할 때는 보상을 잘 주다가 대회에 나가서는 갑자기 강화를 완전히 끊어버리는 것이죠. 음, 그리고 또 하나는 더 오랜 시간 동안 일하게 하는 것과 강화를 줄이는 것을 혼동하는 것이라고 생각해요. 그 둘은 완전히 다른 문제입니다. 강화 줄이기는 개별 행동을 강화하는 것과 관련이 있으며, 모든 행동은 저마다의 간헐적 강화 스케줄을 따르게 됩니다. 더 오랜 시간 동안 작업하는 것은 정신적 지구력에 관한 것으로, 개가 작업 수행 능력이 떨어지기 전까지 얼마나 오랫동안 작업할 수 있는지에 대한 것입니다. 제가 CTP(임계 시점)라고 부르는 것, 즉 작업 수행 능력이 저하되기 시작하는 시점에 도달하기까지의 시간입니다. 집중력이 떨어지기 시작하고, 더 나은 할 일을 찾기 시작할 것입니다. 주변을 둘러보고 바닥 냄새를 맡기 시작하거나, 큐 행동 수행을 멈출 것입니다. 하지만 이는 강화율이나 보상을 줄이는 것과는 관련이 없으며, 그것은 정신적 지구력을 키우는 것, 즉 보상 사이의 작업 시간을 늘릴 수 있는 능력에 관한 것이며, 이는 사실 네, 어느 정도 별개의 문제이며 종종 보상을 줄이는 것과 혼동되기도 합니다. 하지만 실제로는 상당히 별개의 문제이며, 제 생각에는 대회장에 들어가기 전에
environments and they also want to stop using the treats and toys at the same time yeah and you can see why that maybe makes it really hard to accomplish your own goals which is you know where i was going to ask you next which is why do you think so many teams kind of struggle with removing reinforcement sounds like maybe that's part of the answer yeah yeah i think um sometimes thinking about their end goals too much uh before they even start training the behavior they're already worried about like right well i need to really not have this need too much reinforcement you're like no no the dog's learning the behavior at this point they need lots of reinforcement and so that like trying to limit the rewards too soon in the process or not building motivation to perform the behavior enough um not building enough a lengthy enough history of reinforcement before they're quickly trying to drop out the treats and toys um yeah uh maybe they they still have rewards present before cuing the behavior that interferes a lot with the situation so we need to have stopped that happening before we even reduce reinforcement um so reducing rate of reinforcement too quickly is a big reason why people have problems suddenly skipping too many rewards in succession and like you said that's common in trialing they'll they'll keep the rewards up in training then they go to a competition and suddenly drop all the reinforcement in one go um so and i think the other thing is confusing working for longer periods of time with reducing reinforcement they're two very different things reducing reinforcement is specific to reinforcing individual behaviors every behavior is on its own intermittent schedule of reinforcement working for longer periods is about mental stamina that's about how long can a dog work before they reach what i refer to as the ctp the critical time point the point where there's going to be deterioration there's going to start to be some distractibility they're going to start looking for something better to do they're going to look around start sniffing the ground you know stop performing the cue behaviors but that's not to do with the rate of reinforcement or anything reducing reinforcement that's to do with building that mental stamina that ability to work for longer uh between rewards which is actually yeah a little bit of a separate a separate thing often it does get tangled into reducing reinforcement but it is actually quite separate um and i think yeah just not preparing for long periods of work
22:04
간식이나 장난감 없이 긴 시간 작업할 준비를 하지 않는 것이 문제입니다. 그래서 사람들은 그냥 자신의 개가 갑자기 3분 동안 간식이나 장난감 없이 작업하는 것을 눈치채지 못할 것이라고 생각하지만 그렇지 않습니다. 연습할 때 안정적이고 빈번하게 할 수 없다면, 아시다시피 대회장에서 처음으로 시도하는 것은 아마 최선의 방법이 아닐 것입니다. 음, 그리고 제 생각에 다소 관련된 문제로, '운동 종료 큐'를 훈련하지 않는 것이 있습니다. 많은 사람들이 '끝났다'거나 간식을 주는 것과 같은 '세션 종료 큐'는 가지고 있지만, 운동 종료 큐를 훈련하지 않는 것은 정말 까다로울 수 있습니다. 왜냐하면 스포츠에서는 리셋하고 여러 다른 운동을 수행해야 하기 때문입니다. 많은 사람들이 모든 세트나 모든 작업 단계를 보상 마커로 끝내기 때문에, 점점 더 긴 시간 동안 작업하게 되지만 결국 공을 던지거나 터그 놀이를 하거나 간식을 주는 것으로 끝냅니다. 간식을 던지며 끝내곤 하죠. 하지만 대회장에 나갔을 때는 어떻게 하실 건가요? 그 훈련을 끝내고 걸어 나가서 다음 훈련을 준비해야 합니다. 개는 내 보상은 어디 있지? 간식은 어디 있지? 보상은 어디 있지? 라고 생각할 겁니다. 그래서 훈련 종료 신호를 가르쳐서 개가 아, 훈련이 끝났구나, 이제 잠시 머리를 식혀도 되겠네, 우리가 이동해서 다음 훈련을 다시 준비하겠구나 하고 이해하게 하면, 전체 과정을 억지로 계속 진행할 필요가 없습니다. 예를 들어 이번 훈련을 마치고 다음 훈련으로 가기 위해 힐링(나란히 걷기)을 한다고 할 때, 우리는 개가 잠시 머리를 식힐 시간을 갖고, 자유 시간이 조금 주어졌다는 것을 알게 하길 원합니다. 다음 훈련 장소까지 우리와 함께 움직이면 되지만, 보상을 놓쳤다는 느낌을 받지 않으면서도 그것이 가능해야 합니다. 음, 맞아요. 그것도 사람들이 겪는 어려움 중 중요한 이유라고 생각해요. 간식이나 장난감을 주지 않는 확실한 훈련 종료 신호가 없어서 대회장에서 어려움을 겪는 것이죠. 정말 흥미롭네요. 우리가 여기에서 정말 많은 요소를 분해해 보았잖아요. 개별 행동에 대한 보상을 줄이는 것, 그리고 이미 보상을 줄여놓은 여러 가지 행동을 하나의 루틴으로 수행할 수 있도록 정신적 지구력을 기르는 것이 있죠. 그리고 또 다른 부분은, 결국 대부분의 대회 장소나 스포츠에서는 한 가지를 끝내고, 휴식한 뒤, 다음 것으로 넘어가는 것을 요구하는데, 다시 말하지만 고전적인 보상 없이 그 모든 과정들을 전환해야 한다는 점이죠. 그러니까 이건 정말
without treats and toys before entering the competition ring so people just you know sort of think their dog won't notice that they're just suddenly working for you know three minutes without any treats and toys it's like yeah if you can't do it in practice reliably frequently then you know probably not the best thing to just try and do it for the first time in the competition ring um and i think also i guess somewhat related issue um is not training an end of exercise cue so a lot of people have an end of session cue like you know all done do a treats gutter but not training an end of session end of exercise cue can be really tricky um when you're in a sport that requires you to reset and do multiple different exercises because a lot of people end every set every section of work with a reward marker so they work for maybe longer and longer and longer periods but then they end with throwing a ball or end with a strike on a tug or end with giving treats end with a treat toss it's like but what are you going to do when you're in the ring and you just need to end that exercise and then walk off and set up for the new exercise dog's gonna be going where's my reinforcement where's my treats where's my reward so training that end of exercise cue where the dog understands oh end of exercise okay i can have a little mental break we're going to go and reset for another exercise stops us from having to actually just work through the whole thing like i'm going to end this exercise then i'm going to heal to the next exercise you know we want that dog to have that little mental break and know that they're on a little bit of semi free time they can just move with us to the next exercise but they need to be able to do that without feeling like they they missed out on a reward um so yeah i think that's another important point where reason why people have trouble uh in competition rings when they haven't got a solid end of exercise um cue that doesn't involve giving treats and toys that's so interesting i mean we've pulled apart so many different pieces here right there's the reducing reinforcement for the individual behavior there's the building the mental stamina so you can do a whole routine of multiple behaviors that you have already reduced reinforcement for and then there's the piece of like okay yes you you're ultimately most competition venues or sports are going to require finishing one thing break next thing and again you have to transition all of those pieces without um classical reinforcers right so it's just so much more
24:23
독 스포츠를 처음 접하는 사람들이 제대로 이해하기에는 훨씬 더 복잡한 것 같아요. 그래서 우리가 이런 이야기를 나누는 이유 중 일부는, 당신이 8월 일정에 '보상 줄이기'에 대한 새로운 수업을 개설했기 때문이죠. 지금 등록이 가능하니까 사람들이 신청할 수 있어요. 음, 그 수업에서 어떤 부분을 다룰 예정인지 좀 더 공유해 주시겠어요? 수업과 누가 참여하면 좋을지 알려주세요. 네, 기초 교육 섹션에 있습니다. 콘텐츠가 특정 스포츠에 국한되지 않아서 '훈련 발전시키기: 열정은 유지하며 보상 줄이기'라는 제목을 붙였습니다. 제목에 '보상 줄이기'가 포함되어 있지만, 수업에서는 그보다 훨씬 더 많은 내용을 다룹니다. 1주 차에는 트릭과 행동을 만드는 작업을 많이 합니다. 즉, 셰이핑과 루어링에 대해 이야기하고, 신호에 맞춰 행동이 나오도록 합니다. 현재 진행 중인 트릭을 평가하고 어떤 범주에 해당하는지 분류합니다. 아직 습득 단계이거나 초기 행동-결과 단계에 머물러 있는지 확인하고, 아직 루어링과 셰이핑을 많이 해야 하는 단계인지, 신호를 줄 준비가 되었는지 봅니다. 공식적인 신호를 넣을 준비가 되었는지, 아니면 간식이나 장난감이 있을 때만 확실하게 반응하는 단계인지 확인합니다. 트릭 신호를 주기 전에 간식이나 장난감이 있어야 하는 문제를 별도의 과제로 다룹니다. 전반적으로 잘 되더라도 소품을 너무 많이 사용하고 있다면, 해결해야 할 잠재적인 문제로 보고 2주 차에 다루지만 1주 차부터 평가합니다. 트릭은 정말 잘하는데, 환경이 바뀌면 갑자기 다 무너지는 경우도 확인합니다. 이것도 수업 후반부에 다루는 내용이고, '잠깐, 이 트릭에 문제가 있네?'라는 경우도 확인합니다. 골드 학생들은 '이 트릭을 가르쳤는데 왜 안 되지?'라고 생각할 것입니다. 간식이나 장난감이 없으면 트릭이 안 되는 이유를 알 수 없는 상황들을, 항상 정확하게 수행되지 않는 문제일 수도 있다는 점을 살펴보려 합니다. 항상 정식 신호에 맞춰 행동이 나타나는 것은 아니거나 때로는 반복된 신호가 필요할 수도 있어서 우리는 그런 부분들을 살펴볼 예정이며, 그게 1주차 내용이고 2주차는 강화 줄이기와 소품 사용 및 소거에 관한 내용이며 3주차는 생성 및 유지에 관한 것입니다
complex i think than people who are new to dog sports take the time to really understand um so part of the reason we're talking about all this uh is you've got a new class right on reducing reinforcement on the calendar for august it is now open for registration people can go sign up um you want to share a little more about the class kind of which pieces there you're going to cover in the class and maybe who should consider joining you sure um so it's in the foundation foundation section because the content's not specific to one sport it's called progressing your training reducing reinforcement without reducing enthusiasm so even though reducing reinforcement is in the title the class covers like a lot more than just reducing reinforcement so in week one we do a lot of the building of the tricks and the behaviors so we're talking about shaping and luring uh we get the behaviors happening on cue um we're assessing current tricks and sort of saying okay which category do they fall into are they uh still at the point where you know they're they're at that acquisition phase really or the early action outcome phase so we're still doing lots of luring and um shaping they're not ready to put a cue there are they at the next level where we're kind of ready to put a formal cue in there um are they at the point where they actually are reliably happening but only when the treats and toys are present before the tricks queued and so we look at that as a separate uh issue that we work with um we look at is everything going well but we're still using a lot of props so we're talking about um that as being a potential um thing we need to work on we actually work on that in week two but we're assessing those uh exercises in week one we look at uh are we at the point where maybe it's that the trick's really really good but when we go into a different environment suddenly it all falls apart okay that's something we deal with later in the class as well and then we're also saying hang on there's a problem with this trick so the gold students will be saying i've got this trick and i don't know why it's not working but something's going wrong when i don't have the treats and toys the trick's not working and so we're going to be um looking you know maybe it's not always um accurate it's not always occurring on the formal cue or maybe it sometimes needs repeated cues so we're going to look at those sorts of things and that's in our week one week two is all about reducing reinforcement and using and fading props week three is about creating and maintaining
26:45
동기 부여와 보상 이벤트 리허설, 그리고 보상 배치의 효과적인 사용에 관한 내용이고 4주차는 멀리 떨어진 보상을 사용하는 것에 관한 내용입니다. 즉 반려견이 보상을 뒤에 두고 갈등 없이 나가서 작업을 수행한 뒤 다시 돌아올 수 있게 하는 것이죠. 그리고 우리가 말하는 것은 정신적 지구력을 길러 더 오랜 시간 작업하는 것입니다. 이미 간헐적 스케줄로 설정된 모든 행동을 하나로 연결해서 더 길고 긴 시간 동안 작업할 수 있게 하는 것이며, 그러한 이유로 보상 지연에 대해서도 다룹니다. 5주차에는 우리가 이야기했던 운동 종료 신호 훈련과 신호에 대한 예측을 없애는 것에 대해 이야기합니다. 즉 신호를 기다리는 행동을 훈련하고 강화하는 것인데, 이는 많은 반려견에게 중요한 부분입니다. 특히 소품을 사용할 때 소품을 내려놓기만 하면 반려견이 바로 행동을 시작해 버리곤 하죠. 그래서 잠깐, 우리는 신호를 줄 때까지 기다리길 원한다고 말하며 그에 대해 논의합니다. 그리고 간식과 장난감 없이 대회에 나가는 준비에 대해 이야기합니다. 6주차에는 간식과 장난감 주변에서 작업하는 것에 대해 이야기합니다. 우리가 행동을 계속 신뢰할 수 있게 유지하고 근처에 간식 통이 있거나 바닥에 장난감이 있거나 흥미로운 냄새가 나거나 떨어진 음식이 있어도 집중력을 유지하게 하려는 것입니다. 그것이 6주차에 다룰 내용이고요. 네, 꽤 꽉 찬 수업이에요. 아까 누가 들으면 좋을지 물어보셨죠? 음, 제 생각에는, 솔직히 말해서 모두에게 적합할 것 같아요. 제가 대화하는 모든 사람이 제 수업이 잘 맞는다고 말할 것 같거든요. 모두에게 적합할 것이라고 생각합니다. 솔직히 말해서 저는 이 수업이 정말 광범위한 사람들에게 큰 도움이 될 것이라고 생각합니다. 반려견 훈련을 처음 접하는 분들이나 긍정 강화 기반 훈련을 이제 막 시작하는 분들, 그리고 자신의 기술을 더욱 다듬고자 하는 숙련된 훈련사들 모두에게 훌륭하다고 생각합니다. 또한 반려견 훈련 수업이나 스포츠 독 수업을 가르치는 분들에게도 좋을 것입니다. 왜냐하면 이론 주제들이 서로 다른 반려견들에게 왜 우리가 다른 선택을 할 수 있는지에 대한 탄탄한 이해를 제공하기 때문입니다. 음, 숙련된 분들을 위한 이론 부분은 그들이 지적 호기심을 마음껏 충족시킬 수 있게 해줄 것이라고 생각합니다. 정말 멋진 내용들이 많거든요. 종이 자료들을 좀 찾아보자면, 1주 차에는 큐 현저성(cue salience), 오버섀도잉(overshadowing), 큐 전이(cue transfers), 소거(extinction), 미신적 행동(superstitious behaviors) 같은 재미있는 주제들에 대해 다룹니다. 그래서 숙련된 훈련사들을 위한 이론 파트도 아주 멋질 것이라고 생각합니다. 하지만 새로운 훈련사들을 위한 단계별 실습 과정도 준비되어 있습니다.
motivation and rehearsing like reward events and effective use of reward placement week four is about using remotely placed rewards so having our dog being able to leave rewards behind in an unconflicted way and go out and work and then return to them and we're talking then about building mental stamina working for longer periods using all these behaviors that are already on an intermittent schedule but we're piecing them all together and being able to work for longer and longer periods um and we're talking then about delayed rewards for that reason as well week five we're talking about um the training the end of exercise cue that we talked about and eliminating anticipation of cues so training and reinforcing the behavior of waiting for a cue which is a big one for a lot of dogs especially when we're using props we just put the prop down they just start doing the behavior and then we're like hang on we want them to wait until we cue it so we talk about that uh and we talk about preparing for competing without treats and toys and then in week six we just talk about working around treats and toys so when we want to be able to keep those behaviors reliable and occurring in a focused way even when there are treat containers nearby or toys on the ground or interesting smells or drop food so that's what we're talking about in uh week six so yeah it's um pretty jam-packed it's a pretty packed class yeah and i think you did also ask who might it suit um i think it's gonna suit i honestly think i mean everyone i'm sure everyone you talk to would go well my class will suit everyone and but honestly i think this class will be great for a big range of people so people who are brand new to dog training or brand new to positive reinforcement based training as well as experienced trainers looking to further refine their skills um i think it's great for people who instruct probably dog training classes as well or sport dog classes because the theory topics provide a lot of solid understanding of why we might make different choices with different dogs um and i think the theory uh for the experienced people will let them get their geek on because there's lots of really cool you know i think i'm just grabbing for some pieces of paper here but like week one we talk about cue salience and overshadowing and cue transfers and extinction and superstitious behaviors and all these sorts of fun things so uh i think the theory part for the experienced trainers will be really cool uh but there's also just uh the practical stuff is step-by-step
29:12
사실 1주 차 수업에서 저는 훈련을 이제 막 시작하신 분들이라면 1주 차에 포함된 특정 이론 주제 세 가지는 읽지 않고 넘어가도 좋다고 말씀드립니다. 왜냐하면 너무 압도당하지 않기를 바라기 때문입니다. 다른 이론들도 아주 많으니 정말 열의가 있으시다면 나중에 돌아와서 다시 읽어보시면 됩니다. 그 외의 나머지 부분은 훈련을 시작한 지 얼마 안 된 분들이나 현재 진행 중인 트릭 훈련에서 잘 안 풀리는 문제가 있는 분들을 위한 매우 실용적인 단계별 과정입니다. 좋습니다. 마지막으로 하실 말씀이나 청취자들에게 꼭 전달하고 싶은 핵심 포인트가 있을까요? 우리가 무엇을 더 알아야 할까요? 저는 '강화를 줄이기에 너무 늦은 때란 없다'는 점을 다시 강조하고 싶습니다. 많은 분들이 '너무 일찍 줄이고 싶지 않지만, 또 너무 늦게 줄이고 싶지도 않다'고 생각하시는데, 사실 너무 늦은 때란 없습니다. 너무 오래 보상하는 것이 차라리 부족한 것보다 훨씬 낫거든요. 보상을 너무 일찍 또는 너무 빠르게 줄이려고 하는 것이죠. 핵심은 연속 강화 스케줄에서 간헐적 강화 스케줄로 점진적으로 전환하는 것입니다. 그리고 각 행동을 개별적으로 다루어야 합니다. 왜냐하면 일부 행동은 다른 행동보다 훨씬 더 높은 강화 비율로 유지되어야 하기 때문입니다. 그러니 각 행동을 별도로 다루고, 결코 너무 늦지 않았다는 점을 기억하세요. 설령 지금까지 모든 행동에 대해 항상 간식을 너무 많이 줘 왔고, '아차, 내가 모든 것에 항상 간식을 주는 끔찍한 습관을 들였구나'라고 생각하신다면 강의를 신청하세요. 왜냐하면 그 상황을 바꾸는 데에는 결코 너무 늦지 않았기 때문입니다. 그 행동들이 정말 확실하고 신속하고 정확하게 수행되도록, 매번 간식이나 장난감이 없어도 되게끔 바꿀 수 있습니다. 멋지네요. 네, 오늘 이 모든 것에 대해 이야기해 주러 팟캐스트에 출연해 주셔서 정말 감사합니다. 아닙니다. 다시 대화할 수 있어서 정말 즐거웠고, 강의가 시작되는 것도 기대하고 있습니다. 그럼요. 청취해 주신 모든 분께 감사드리며, 다음 주에 다시 돌아오겠습니다. 놓치지 마세요. 아직 구독하지 않으셨다면 아이튠즈나 사용하시는 팟캐스트 앱에서 저희 팟캐스트를 구독하셔서 다음 에피소드가 나오는 즉시 휴대폰으로 자동으로 다운로드되도록 하세요. 오늘의 방송은 펜지 도그 스포츠 아카데미(Fenzy Dog Sports Academy)에서 제공합니다. 데니스 펜지에게 특별히 감사드립니다. 이 팟캐스트를 지원해 주셔서 감사합니다. 음악은 bensound.com에서 로열티 프리로 제공되었습니다. 여기에 사용된 곡 제목은 바디(Body)입니다. 오디오 편집은 크리스 랭이 제공했습니다. 다시 한번 청취해 주셔서 감사드리며, 즐거운 훈련 되세요. 엄마
processes for the newer trainers and in fact i think in week one i even say if you're brand new to training maybe skip reading these you know certain three topics that are in week one um because you know i don't want you to get overwhelmed there's tons of other theory and if you're really keen you can go back and reread those um but the rest of it is very much practical step-by-step processes for um yeah people that are newer to training or people that are having current issues with something to do with their tricks that aren't working well all right final thoughts key points you want to leave folks with what else do we need to know on this um i think i just hop back to that it's never too late to reduce reinforcement i think a lot of people go oh i don't want to do it too soon but i don't want to do it too late and it's like well it's never too late like rewarding for too long is way better than trying to reduce the rewards too early or too quickly so the key is gradual shifting though from that continuous schedule to the intermittent schedule uh and treating each behavior in isolation so some behaviors will need to be maintained on that much higher rate of reinforcement than other behaviors so just yeah treating each behavior separately and just remembering it's yeah never too late to even if you've been just giving lots of treats all the time for everything and you're like oh no i've got in this horrible habit if i give treats all the time for everything um then come and join the class because it is never too late to change that uh to the situation where those behaviors are performed really reliably really rapidly and accurately without needing uh treats and toys after every performance of the behavior awesome all right well thank you so much for coming on the podcast to talk about all this no worries it's been fantastic talking to you again and yeah i'm looking forward to the class starting absolutely and thanks to all of our listeners for tuning in we'll be back next week don't miss it if you haven't already subscribe to our podcast in itunes or the podcast app of your choice to have our next episode automatically downloaded to your phone as soon as it becomes available today's show is brought to you by the fenzy dog sports academy special thanks to denise fenzy for supporting this podcast music provided royalty free by bensound.com the track featured here is called body audio editing provided by chris lang thanks again for tuning in and happy training mom