<![CDATA[Episode 64: Variable Reinforcement is a LIE]]>
Ivan Balabanov<![CDATA[Training Without Conflict® | Dog Training Podcast]]> · 팟캐스트 · 21분
반려견 훈련에서 흔히 활용되는 간헐적 강화 계획이 행동 유지에 필수적이라는 통념은 지나치게 단순화된 접근 방식이다. 행동 습득 이후에도 연속 강화 계획을 지속적으로 사용하는 방식은 반려견의 동작을 더 날카롭고 정확하게 개선하는 지렛대 역할을 수행한다. 반려견이 보상을 기대하는 상황에서 갑작스럽게 보상을 제한하면 발생하는 소거 폭발 현상은 단순히 부정적인 장애물이 아니라, 반려견의 더 높은 집중과 과장된 노력을 이끌어내어 훈련 기준을 높이는 강력한 교육적 도구가 된다. 짖음이나 입질과 같은 좌절의 표현 또한 훈련사와 반려견 간의 강화 갈등에서 기인하는 피드백으로 해석되어야 하며, 이를 적절히 다룰 경우 오히려 생산적인 훈련 성과로 전환된다. 진정한 행동의 지속성은 단순히 무작위 보상을 통한 도박 심리에 의존하는 것이 아니라, 반려견이 훈련 과정 자체를 습관이자 즐거움으로 인식하고 스스로 행동을 강화하는 상태에서 완성된다.
에피소드 64: 가변 강화는 거짓말입니다
<![CDATA[Episode 64: Variable Reinforcement is a LIE]]>
0:00
. 좋습니다. 가변 강화 계획에 대해 여러분이 알고 있는 모든 것이 절반의 진실에 불과하다면 어떨까요? 오늘은 제가 수강생들과 다루는 주제들을 살짝 맛보여 드리고 여러분도 함께 고민해 볼 거리를 드리고자 합니다. 강화 계획은 반려견 훈련과 심리학에서 매우 흔하게 언급되는 개념입니다. 훈련사들이 이렇게 말하는 것을 분명 들어보셨을 겁니다. 무작위 보상이 행동을 더 확실하게 강화하니까 가변 강화 계획으로 전환하세요. 그러고 나서 슬롯머신 예시를 들겠죠. 사실 많은 분들이 이미 그렇게 하고 계실 겁니다. 물론 그 말에도 일리는 있습니다. 하지만 거기서 멈춘다면 아마 더 큰 그림을 놓치고 계신 걸지도 모릅니다. 제가 훈련에서 강화 계획을 사용하는 방식은 슬롯머신 논리를 조금 넘어섭니다. 자, 그럼 간단하게 설명해 보겠습니다. 먼저 기초부터 시작하죠. 행동 심리학에는 여러 가지 강화 계획이 있습니다. 연속 강화가 있는데, 기본적으로 모든 올바른 반응마다 보상을 제공하는 것입니다. 새로운 행동을 가르칠 때 가장 좋은 선택이라는 조언이 항상 주어지는데, 이는 사실입니다. 고정 비율 계획도 있습니다. 이는 정해진 횟수의 반응 후에 보상을 제공하는 방식입니다. 예를 들어, IGP의 '짖고 버티기(bark and hold)'를 들 수 있겠네요. 만약 훈련사가 세 번 짖을 때마다 보상한다면, 강아지는 계속 짖기보다는 세 번 짖고 보상을 기다리게 될 가능성이 높습니다. 그것이 바로 고정 비율 계획의 예입니다. 그리고 가변 비율 계획, 즉 무작위 강화 계획이라고도 불리는 방식이 있습니다. 이는 예상치 못한 횟수의 반응 후에 보상을 주는 방식입니다. 라스베이거스 비유가 나오는 지점이 바로 여기죠. 때로는 두 번의 반응 후에, 때로는 다섯 번의 반응 후에 보상을 주는 식으로 계속됩니다. 또 다른 하나는 고정 간격 강화입니다. 일정 시간이 지난 후 첫 번째 반응에 대해 보상합니다. 예를 들어, 개가 30초가 지난 후에 앉았을 때만 강화가 이루어지는 식이죠. 변동 간격 강화입니다. 변동 간격 강화죠. 예측할 수 없는 시간 간격이 지난 후 첫 번째 반응에 대해 보상합니다. 그러니까 때로는 10초 후가 될 수도 있고, 때로는 1분 후가 될 수도 있습니다. 그게 강화 스케줄 매뉴얼의 거의 전부입니다. 자, 반려견 훈련의 실제 현장에서는 대부분 의도적으로 이런 것들을 사용하지 않습니다. 훈련사들은 일반적으로 가르칠 때는 연속 강화 방식을 고수하고, 행동을 유지하려고 할 때는 변동 강화 방식을 사용합니다. 그 지점에서 항상 슬롯머신 비유가 나오죠, 맞나요?
. All right. What if everything you've been told about variable reinforcement schedule is only half the story? Today I want to give you a taste of the kind of topics I go over with my students and hopefully give you something to think about. Schedule of reinforcement is something that gets thrown around a lot in dog training and in psychology. I'm sure you've heard trainers say, switch to variable reinforcement schedule because random rewards make the behavior stick. And then they're going to give you the slot machine example. In fact, probably many of you are doing the same. And sure, there is truth to that. But if you stop there, you are probably missing the bigger picture. The way I use reinforcement schedules in training goes a little bit beyond slot machine logic. So, let me break this down quickly. But first let's start with the basics. In behavior of psychology, there are several reinforcement schedules. We have continuous reinforcement, which basically every single correct response is rewarded. And the advice that's always given, and it's true, is that it's the best option for teaching new behaviors. We also have fixed ratio. That's where the reward is offered after set number of responses. For example, like in, let's say, IGP, the bark and hold. If the trainer rewards every third bark, then the dog most likely is going to bark three times and wait for the reward instead of continuous barking. That's your fixed ratio example. And we have variable ratio, also known as the random schedule of reinforcement. This is where reward, we reward after unexpected number of responses. That's where the Las Vegas analogies come. Sometimes we pay after two, sometimes after five responses, and so on. Another one, fixed interval. We reward for the first response after a set period of time. For example, the dog only gets reinforced if it sits after 30 seconds have passed. Variable interval. Variable interval. We reward for the first response after unpredictable time intervals. So sometimes it can be after 10 seconds, sometimes after one minute. That's pretty much the full manual of reinforcement schedules. Now, the reality in dog training, we don't use most of this deliberately. Trainers typically stick with the continuous reinforcement when teaching and variable reinforcement when trying to maintain the behavior. That's where the slot machine analogy always comes in, right?
3:59
하지만 주의해야 할 다른 점도 있습니다. 때때로 훈련사들은 자신도 모르게 고정 패턴이나 간격 패턴에 빠지곤 합니다. 어쩌면 그들은 항상 정확히 세 걸음의 힐링 후에 보상하거나, 1분간의 엎드려 기다려가 끝날 때마다 보상할지도 모르죠. 훈련사가 의도하지 않았더라도 개는 그런 패턴을 알아차립니다. 그래서 네, 이론적으로는 모든 스케줄이 존재하지만, 반려견 훈련의 실전에서 훈련사들은 연속 강화 스케줄과 변동 강화 스케줄에 크게 의존합니다. 그리고 대화의 대부분이 바로 그 지점에서 막히게 되죠. 이것이 많은 훈련사가 연속 강화로 시작해서 개가 행동을 익히면 변동 강화로 전환하는 이유입니다. 물론 여기에는 과학적 근거가 있습니다. 스키너도 있고, 소거 저항 연구 등 여러 가지가 있죠. 그리고 네, 변동 강화 스케줄로 훈련된 행동에 대한 보상을 완전히 중단하면, 일반적으로 연속 강화로만 훈련된 행동보다 더 오래 지속됩니다. 그러니 이론적으로는 완전히 말이 됩니다. 하지만 슬롯머신 예시를 더 자세히 들여다봅시다. 여기 우리가 종종 잊거나 생각하지 못하는 부분이 있습니다. 대부분의 사람들은 슬롯머신에 중독되지 않습니다. 수백만 명이 라스베이거스를 방문해 재미로 조금씩 즐기다가 그냥 돌아옵니다. 저도 그 좋은 예입니다. 특정 유형의 사람들만이 도박 중독의 굴레에 빠집니다. 왜일까요? 가변 강화 계획이 마법처럼 모두를 끌어들이지는 않기 때문입니다. 중독은 그 계획이 개인, 성격, 뇌 화학, 취약성 등과 상호작용할 때 발생합니다. 개들도 마찬가지입니다. 모든 개가 예측 불가능성에 동기부여를 받는 것은 아닙니다. 많은 경우 가변 강화 계획이 반드시 집착을 만드는 것은 아닙니다. 그저 하나의 패턴일 뿐이죠. 그래서 슬롯머신 비유는 다소 지나친 단순화입니다. 끈기는 단순히 무작위성 때문에 생기는 것이 아닙니다. 정말로 중요한 것은 행동 자체가 즐겁고 의미 있으며 습관이 되는 시점입니다. 그것이 행동을 완벽하게 만드는 요소입니다. 하지만 진짜 문제는 여기에 있습니다. 개들은 스키너 상자 속의 비둘기가 아니며, 슬롯머신 레버를 당기는 도박꾼은 더더욱 아닙니다. 제 생각에 그런 논리는 실험실 수준에서 멈춰야 합니다. 훈련에서 행동은 단순히 무작위성 때문에 유지되지 않습니다. 습관이 되고, 개가 실제로 그 행동을 즐기기 때문에 유지되는 것입니다. 행동이 진정한 습관으로 자리 잡으면 소거는 사실 큰 위협이 되지 않습니다. 그리고 행동 자체가 스스로 강화될 때, 즉 개가 작업 자체에서 기쁨을 느낄 때 그 어떤 슬롯머신도 이를 따라올 수 없습니다. 제 훈련 방식에서 두 가지 짧은 예시를 들어보겠습니다. 힐링(나란히 걷기)을 예로 들어보죠. 힐링을 처음 가르칠 때는 당연히 연속 강화 계획을 사용합니다. 그래서 모든 올바른 단계마다 보상이 주어집니다.
But there is also something else to watch for. Sometimes trainers fell into fixed or interval patterns without ever realizing it. Maybe they always reward after exactly three steps of healing or always at the end of one minute down stay. The dog notices those patterns even if the trainer doesn't intend them. So, yes, technically all the schedule exists, but in dog training, in practice, dog trainers lean heavily on continuous and variable reinforcement schedule. And that's where most of the conversation really gets stuck. This is why many trainers start on continuous reinforcement and then switch to variable once the dog knows it. Sure, there is science behind this. You have skinner, you have resistant to extinction research, and the whole thing. And yes, if you stop rewarding completely a behavior that is trained on variable reinforcement schedule, it will usually last longer than one trained only on continuous. So, in theory, it makes total sense. But, let's look closer at that slot machine example. Here is something we kind of forget or don't think of. Most people are not hooked on slot machines. Millions go to Vegas, play a little for fun, and walk away. I'm a prime example of this. Only a certain type of person gets trapped in the cycle of gambling addiction. Why? Because variable schedules don't magically hook everyone. Addiction happens when the schedule interacts with the individual, their personality, the brain chemistry, susceptibility, and so on. And dogs are no different. Not every dog is motivated by unpredictability. For many, variable schedule of reinforcement doesn't necessarily create obsession. It's just a pattern. So, the slot machine analogy is kind of oversimplification. Persistent isn't just about randomness. What really matters is when the behavior itself becomes enjoyable, meaningful, and habitual. That's what makes it bulletproof. But, here is the real problem. Dogs aren't pigeons in Skinner boxes, and they're definitely not gamblers pulling levers at the slot machine. So, that kind of logic stops in the laboratory to me. In training, behaviors don't just survive because of randomness. They survive because they become habits, and because the dog actually loves doing them. Once a behavior is a true habit, extinction isn't really a threat. And, when a behavior is self-reinforcing, when the dog takes joy in the work itself, no slot machine can compete with that. I'll give you two quick examples from my own training. Let's take healing. When I start teaching healing, of course, I use continuous schedule of reinforcement. So, every correct step gets paid.
7:58
그 명확함이 필수적이죠, 그렇죠? 하지만 시간이 지나면 무언가가 변합니다. 제 개들에게 힐링은 단순한 보상을 넘어선 것이 됩니다. 그것은 습관이자 즐거움이 됩니다. 녀석들은 도전과 정확성, 그리고 저와의 상호작용을 갈망하게 되죠. 그 지점에 이르면, 힐링을 유지해 주는 것은 더 이상 간헐적 강화 계획이 아닙니다. 실제로는 개가 느끼는 즐거움 그 자체가 그것을 유지해 줍니다. 이제, 센드 어웨이(send away)를 예로 들어보죠. 더 많은 거리, 더 빠른 속도, 더 날카로운 정확성을 원한다면, 이번에도 무작위 강화 계획은 저를 그 목표에 도달하게 해주지 못할 것입니다. 이때야말로 연속 강화 계획이 다시 빛을 발하는 순간입니다. 저는 모든 개선 사항과 더 날카로워진 노력 하나하나에 보상을 줍니다. 그렇게 행동에 대한 기준을 높여가는 것입니다. 그러니 이렇게 생각하시면 됩니다. 간헐적 강화 계획은 행동을 현재 상태로 고착시키는 경향이 있습니다. 연속 강화 계획은 저로 하여금 그것을 더 날카롭게 다듬을 수 있게 해줍니다. 자, 여기 강화 계획이 일반적으로 가르쳐지는 방식과는 완전히 상반되는 사실이 하나 있습니다. 대부분의 훈련사와 교재들은 여러분이 원한다면 행동을 가르치는 습득 단계에서는 연속 강화 계획을 사용하고, 그 후에는 행동을 더 지속시키기 위해 빠르게 간헐적 강화 계획으로 넘어가라고 말할 것입니다. 그리고 소거에 저항하도록 만드는 것이죠. 그리고 유지 단계에 대해서는 당연히 그러한 이유 때문에 간헐적 강화 계획을 유지하라고 조언할 것입니다. 그것이 여러분이 아는 일반적인 통념입니다. 하지만 제가 더 효과적이라고 발견한 것은 이것입니다. 저는 유지 단계에서도 종종 연속 강화 계획을 계속 사용합니다. 그리고 제가 '고유지 행동'이라고 부르는 것들에 대해 그렇게 합니다. 그게 무슨 뜻이냐고요? 앉아(sit)를 예로 들어보겠습니다. 대부분의 개에게 앉기는 자연스럽고 쉬운 반응이며, 배우기 쉬운 동작입니다. 일단 훈련이 되면, 매번 제대로 된 앉기 자세를 얻을 수 있습니다. 하지만 대회에서는 '제대로 된' 앉기만으로는 충분하지 않습니다. 우리는 더 날렵하고, 더 빠르고, 더 정확한 동작을 원합니다. 개의 본능적인 '원래 이렇게 하는 거예요' 식의 동작보다 더 나은 무언가가 필요합니다. 그래서 그 수준을 한 단계 더 높이기 위해, 저는 지속적인 강화 계획(continuous schedule of reinforcement)에 따라 계속 보상을 줍니다. 왜일까요? 개의 기본 행동 이상의 것을 요구할 때는, 그 날렵함을 유지하기 위해 지속적인 강화가 필요하기 때문입니다. 물론, 여기에는 주의해야 할 점이 있습니다. 조금씩만 더 요구할 수 있다는 점입니다. 너무 무리하게 밀어붙여서 계속해서 더 많은 것을 요구하다 보면, 결국 한계점에 다다르게 됩니다. 그렇게 되면 행동이 개선되기는커녕 오히려 무너지기 시작할 것입니다. 개는 말 그대로 자신의 한계 이상으로 행동할 수 없기 때문입니다. 여기 또 다른 점이 있습니다. 트레이너들은 더 나은 시도에만 간헐적 강화(variable schedule)로 보상하면 기준을 높일 수 있다고 말하곤 합니다.
That clarity is essential, right? But, over time, something changes. For my dogs, healing becomes more than just a paycheck. It becomes a habit and a joy. They crave the challenge, the precision, the interaction with me. At that point, variable schedule of reinforcement isn't what holds healing together. It is actually the dog's own enjoyment. Now, let's take the send away. If I want more distance, more speed, sharper precision, then, again, random reinforcement schedule is not going to get me there. This is where continuous schedule of reinforcement shines again. I reward every improvement, every sharper effort. That's how I will raise the criteria for the behavior. So, you can think of this way. Variable schedule of reinforcement tends to freeze the behavior where it is. Continuous schedule lets me sharpen it. Now, here is something that goes completely against the way reinforcement schedules are usually taught. Most trainers, and most textbooks, if you want, will tell you, use continuous schedule of reinforcement during the acquisition state to teach the behavior, then quickly move to variable schedule of reinforcement to make the behavior more persistent. And resistant and resistant to extinction. And for the maintenance stage, of course, they would advise to stay with variable schedule of reinforcement because of that. That's your conventional wisdom. But here is what I found works better. I often continue to use continuous schedule of reinforcement during the maintenance stage. And I do it for what I call high maintenance behaviors. What do I mean? Let's take the sit as an example. For most dogs, sitting is a natural and easy response, easy thing to learn. Once trained, we'll get a decent sit every time. But in competition, a decent sit isn't enough. We want something sharper, something faster, more precise. Something better than the dog's natural, this is how I do it. And so to get that extra level, I keep rewarding on continuous schedule of reinforcement. Why? Because when I am asking for more than the dog's default behavior, it needs constant reinforcement to maintain that sharpness. Now, of course, there is a warning that goes with that. Like you can ask only for a little more. If we push too far, if we demand more and more and more, eventually we're going to hit a breaking point. And instead of improving the behavior, it's going to start to crumble. Because the dog literally cannot perform better than its limit. And here is another point. Trainers will say, just reward the better attempts on a variable schedule and you will be able to raise the criteria.
11:54
하지만 실제로는 그렇게 작동하지 않습니다. 이미 간헐적 강화 계획에 익숙해진 개에게 보상을 주지 않는 것은 전혀 특별한 일이 아닙니다. 개는 실망해서 더 열심히 하거나 더 자주 시도하지 않습니다. 그저 '다음번엔 늘 그랬듯 보상을 받겠지'라고 생각할 뿐입니다. 그것이 제가 높은 수준의 유지가 필요한 행동에 대해 지속적 강화 계획을 고수하는 이유입니다. 이해가 되셨기를 바랍니다. 다시 말하지만, 무작위 강화 계획은 행동을 활기차게 유지해 줄 수 있습니다. 하지만 행동을 더 나은 수준으로 만드는 것은 바로 지속적 강화 계획입니다. 자, 이제 많은 사람이 이야기하지 않는 부분입니다. 제게는 이 부분이 고급 훈련의 진정한 핵심입니다. 우리는 보통 소거 폭발(extinction burst)에 대해 부정적인 방식으로만 듣습니다. 개가 문 앞에서 짖을 때 우리가 강화를 중단하면, 짖는 행동이 완전히 사라지기 전까지 한동안은 오히려 더 심해집니다. 그래서 훈련사들은 흔히 '소거 폭발'이라고 말하며, 그냥 대비하고 있으라고 하죠. 하지만 소거 폭발은 단지 장애물이 아닙니다. 이는 행동을 개선하는 가장 강력한 도구 중 하나가 될 수 있습니다. 그 방법을 알려드리겠습니다. 제가 연속 강화 스케줄을 사용할 때, 개는 매번 보상을 기대합니다. 그러니 갑자기 보상을 주지 않으면 개는 어떻게 할까요? 항의를 하겠죠. 마치 '어이, 내 돈은 어디 갔어?'라고 말하는 것처럼요. 그 항의하는 과정에서 개는 행동을 더 과장되게 표현합니다. 그 과장된 행동은 보물과 같습니다. 그것이 바로 제가 힐을 정교하게 만들고, 자세를 바로잡고, 앉아 훈련 등을 하는 방식입니다. 소거 폭발은 개가 더 많은 것을 시도하도록 유도하고, 저는 그때 보상을 줍니다. 그렇게 기준을 높여가는 것입니다. 이제 이미 가변 강화 스케줄에 익숙해진 개에게 이 방식을 시도한다고 상상해 보세요. 그건 효과가 없을 겁니다. 개는 항의하지 않거든요. 그저 '음, 다음번에 보상을 받겠지'라고 생각할 뿐입니다. 폭발(항의)할 이유가 없는 거죠. 추가적인 노력을 기울일 이유가 없습니다. 그러니 네, 가변 강화는 도움이 되고 행동을 유지하게 해줍니다. 하지만 행동을 더 좋게 만들기 위한 지렛대 역할을 하는 것은 바로 연속 강화 스케줄입니다. 그리고 이것은 대부분의 훈련사가 완전히 놓치고 있는 아주 큰 차이점입니다. 마무리하기 전에, 훈련 과정에서 나타나는 이와 관련된 한 가지 현상을 짚어보겠습니다. 가끔 개들은 훈련 중에 짖거나, 낑낑거리거나, 심지어 핸들러를 살짝 물기도 합니다. 많은 이들이 이를 '리킹(leaking)'이라고 부릅니다. 물론 어떤 이들은 이를 '드라이브'라고 부르기도 합니다. 또 다른 사람들은 이를 불복종과 혼동하기도 합니다. 하지만 대부분의 경우, 그것은 두 가지 이유로 귀결됩니다. 무엇을 기대하는지에 대한 명확성이 부족하거나 강화 일정을 잘못 다루는 경우입니다. 개는 무엇인가를 해야 한다는 요구는 알고 있지만, 정확히 무엇을 해야 할지 완전히 이해하지 못하면 당연히 좌절감이 밖으로 표출됩니다.
But it doesn't work that way. If the dog is already on variable reinforcement schedule, withholding the reward is nothing unusual to them. The dog isn't going to be disappointed and try harder or often more. It just assumes, maybe next time I will get paid as I always do. That's why I stay with a continuous schedule for reinforcement for high maintenance behaviors. Hope that makes sense. So, again, random schedules, we can say, keeps things alive. But it is the continuous schedule of reinforcement that makes behaviors better. Now, here is the part that not many talk about. And for me, it's really the essence of advanced training. We usually hear about extinction bursts in a negative way. A dog barks at the door, we stop reinforcing it, and for a while the barking gets worse before it fades out. So, a trainer is going to say, yeah, that's a typical extinction burst, so just be ready for it. However, extinction bursts aren't just an obstacle. They can be one of the most powerful tools for improving behavior. Here is how. When I'm using a continuous reinforcement schedule, my dog expects a reward every single time. So, if I suddenly hold back, the dog is going to what? Protest. I say, hey, where is my money? And in that protest, the dog exaggerates the behavior. That exaggeration is gold. That's how I sharpen the heel, fasten the waist, sit, and so on. The extinction burst pushes the dog to offer more, and then I reward that. That's how I raise the criteria. Now, picture trying this on a dog already on variable schedule of reinforcement. It's not going to work. The dog doesn't protest. It just assumes, okay, maybe I get paid next time. There is no reason for a burst. There is no reason to make an extra effort. So, yes, variable reinforcement helps, keeps the behavior alive. But it's the continuous reinforcement schedule that gives me the leverage to make behaviors better. And that's a huge difference most trainers completely miss. Before I wrap up, let me touch on something related that shows up in training. Sometimes, dogs will bark, whine, or even nip at the handler during performances. Many call this leaking. Some, of course, call it drive. Others confuse it with disobedience. But most of the time, it comes down to two things. Lack of clarity about what's expected or mishandling reinforcement schedules. If the dog knows there is a demand to do something, but doesn't fully understand what, the frustration, of course, leaks out.
15:42
만약 강화가 잘못 다루어졌다면, 예를 들어 연속 강화 일정에서 너무 빨리 벗어난 경우, 방금 말했듯이 항의는 짖거나 무는 행동으로 나타납니다. 자, 여기서 핵심은 이것입니다. 그 항의가 항상 나쁜 것은 아닙니다. 소거 폭발과 마찬가지로, 그것은 양날의 검입니다. 그것을 인식한다면, 당신은 그것을 활용할 수 있습니다. 그 짖음, 입질, 작은 좌절감의 폭발은 더 날카로운 노력, 더 빠른 속도, 혹은 더 높은 강도 등 당신이 추구하는 무엇으로든 전환될 수 있습니다. 하지만 불행하게도 모든 사람이 이를 이런 식으로 보지는 않습니다. 항상 개가 소리를 낼 때마다 도미넌트 칼라로 개의 숨통을 조이라고 권장하는 소셜 미디어 인플루언서가 떠오릅니다. 그리고 슬픈 점은 훈련사들이 좋아요 버튼을 누르며 이것이 훌륭하다고 생각한다는 것입니다. 그렇지 않습니다. 그것은 끔찍한 조언입니다. 훈련 중에 짖거나 낑낑거리는 것은 개를 질식시키는 것으로 해결되지 않기 때문입니다. 그것은 양질의 능숙한 강화 처리를 통해 해결됩니다. 개는 잘못된 행동을 하는 것이 아닙니다. 개는 당신에게 피드백을 주고 있는 것입니다. 그리고 그 피드백을 제대로 읽는다면, 당신은 그것을 갈등 대신 탁월함으로 바꿀 수 있습니다. 더 큰 그림은 이렇습니다. 변동 강화 일정은 다시 한번 강조하지만 매우 유용합니다. 하지만 그것들이 만능 해결책은 아닙니다. 기준을 높이고 정밀도를 향상하기 위해서는 연속 강화 일정이 매우 중요합니다. 소거 폭발은 단순히 두려워해야 할 대상이 아닙니다. 그것들은 행동을 날카롭게 만드는 최고의 도구 중 하나입니다. 짖음, 낑낑거림, 깨무는 행동 등은 종종 강화 갈등에서 비롯된 항의인데, 무엇을 해야 할지 알면 생산적인 방향으로 이끌 수 있습니다. 라스베이거스에 비유하자면, 네, 앞서 말했듯이 그것은 이야기의 절반일 뿐입니다. 끈기, 끈기는 단지 무작위성에 관한 것이 아닙니다. 진정한 지속성은 상호작용 그 자체에 대한 습관과 즐거움에서 나옵니다. 목표는 우연히 행동이 유지되도록 하는 것이 아닙니다. 목표는 행동을 매우 명확하고, 강력하고, 즐겁게 만들어 습관이 되고, 스스로 강화되며, 결과적으로 완벽하게 만드는 것입니다. 반려견을 훈련하고 있다면 보상을 무작위로 만드는 것에 너무 집착하지 마세요. 스스로에게 물어보세요. 지속적 강화 일정 동안 행동을 아주 명확하게 가르쳤는가? 단순히 상황을 유지하는 것이 아니라 품질을 높이기 위해 강화를 사용하고 있는가? 반려견이 실제로 그 훈련을 즐기는가? 물론입니다. 상호작용을 습관이 될 정도로 충분히 즐기는가? 그 부분에 집중한다면, 소거(extinction)는 훈련사들이 생각하는 것만큼 큰 문제가 아닙니다. 이것은 제가 운영하는 '갈등 없는 훈련을 위한 반려견 훈련사 학교(School for Dog Trainers)'에서 가르치는 여러 흥미로운 내용 중 하나입니다. 물론 강화 일정도 다루지만, 교과서적인 내용을 넘어 제가 '명백한 콘텐츠'라고 부르는 것 그 이상을 가르칩니다.
If reinforcement has been mishandled, let's say, coming off continuous schedule of reinforcement too quickly, the protest shows up as vocalizing or biting, as I just said. So, here is the key. That protest isn't always bad. Just like extinction bursts, it's a two-edged sword. If you recognize it, you can harness it. That bark, the nip, the little explosion of frustration can be channeled into sharper effort, more speed or more intensity, whatever you're after. But, unfortunately, not everyone sees it this way. There is a social media influencer that comes to mind who always recommends stopping dog's air supply with a dominant collar every time it vocalizes. And the sad part is trainers are hitting the like button and thinking this is brilliant. It's not. It's a horrible advice. Because barking or whining in training isn't fixed by choking out the dog. It's fixed by quality and skillful reinforcement handling. The dog isn't misbehaving. It's giving you feedback. And if you read that feedback properly, you can turn it into brilliance instead of a conflict. Here is the bigger picture. Variable schedule, once again, very useful. But they are not the magic bullet. Continuous schedule of reinforcement is critical for raising criteria and improving precision. Extinction births aren't just something to fear. They are one of the best tools for sharpening behavior. The barking, whining, nipping, and so on are often protests born of reinforcement conflict, which can be channeled productively if you know what you're doing. As far as the Las Vegas comparison, yeah, as I said, it's only half the story. Persistent, persistence isn't just about randomness. Real durability comes from habit and joy of the interaction itself. The goal isn't to keep behaviors alive by chance. It's to make them so clear, so strong, and so enjoyable that they become a habit, self-reinforcing, and therefore bulletproof. If you're training your dog, don't obsess over making rewards random. Ask yourself, have I made the behavior crystal clear during the continuous schedule of reinforcement? Am I using reinforcement to raise quality, not just keep things going? Does my dog actually enjoy the work? Of course. Does it enjoy the interaction enough that it becomes a habit? If you focus there, extinction isn't the problem trainers make it out to be. So, this is just one of the many interesting things that I teach at my School for Dog Trainers, training without conflict. Of course, we cover reinforcement schedules, but we do go beyond the textbook and beyond what I call obvious content.
19:39
우리는 소거 폭발(extinction burst)을 도구로 사용하는 법, 기준을 높이는 법, 발성 및 갈등 행동을 해석하는 법을 가르칩니다. 그리고 스스로 강화되는 행동을 만드는 법도 가르칩니다. 진정한 훈련은 단지 슬롯머신에 관한 것만이 아니기 때문입니다. 그것은 명확성, 즐거움, 정밀함, 그리고 유대감에 관한 것입니다. 들어주셔서 감사합니다. 여러분의 생각을 알려주세요. 잘 모르겠습니다.
We teach how to use extinction bursts as a tool, how to raise criteria, how to interpret vocalization and conflict behaviors, and how to create actions that become self-reinforcing. Because real training isn't only about slot machines. It's about clarity, joy, precision, connection. Thanks for listening. Let me know what you think. I'm not sure.