Dog Training Reinforcement Process: How, When and Where Are You Rewarding Your Dog?
Susan GarrettDogs That · 영상 · 9분
반려견 훈련 강화 과정: 반려견에게 어떻게, 언제, 어디서 보상하고 계신가요?
Dog Training Reinforcement Process: How, When and Where Are You Rewarding Your Dog?
0:00
강화 기반 프로그램으로 훈련할 때에 대해 이야기해 봅시다. 즉, 신체적 교정이나 언어적 위협 없이 반려견을 훈련하기로 결정했다는 뜻이죠. 안 돼, 아니야 같은 말을 하고 싶지 않거나, 목줄을 잡아당기는 행동을 원하지 않는 경우입니다. 그렇게 훈련하기로 마음먹었다면 가장 중요한 도구는 바로 강화입니다. 그래서, 강화를 사용하는 방법과 이해도에 있어 매우 능숙해져야 합니다. 무엇을 알아야 할까요? 우선 반려견에게 가장 좋은 강화물이 무엇인지 알아야 합니다. 첫 번째로, 기대치를 훨씬 뛰어넘고, 아주 열광적이며, 세상에, 말도 안 돼, 이 강화물이 나오면 정신을 못 차릴 정도로 좋아하는 것이 무엇인지 알아야 합니다. 또한 무엇이 정말 좋은 강화물인지도 알아야 하죠. 그리고 무엇이 적절한 수준의 강화물인지도 알아야 합니다. 이러한 것들은 반려견의 연령, 경험, 훈련 수준 그리고 훈련 장소에 따라 달라질 것입니다. 즉, 지금 이 환경과 이 훈련 단계에서 무엇이 최고이고, 무엇이 좋으며, 무엇이 적절한지를 알아야 합니다. 그렇다면 방해 요소가 있는 상황에서는 어떨까요? 환경적인 방해 요소가 많다면, 더 높은 수준의 강화를 사용해야 하기 때문입니다. 따라서 이전에는 적절했던 것이 주변에 방해 요소가 많을 때는 적절하지 않게 될 것입니다. 반려견을 훈련할 때마다 어디서 훈련하는지, 그리고 무엇이 최고이고, 무엇이 좋으며, 무엇이 적절한지 고려해야 합니다. 왜냐하면 그것들은 계속 변할 것이기 때문입니다. 또한 훈련하는 환경에 의해 어떤 제약이 있는지 알아야 합니다. 예를 들어, 훈련 수업 중이라면 소리 나는 장난감을 사용할 수 없습니다. 다른 모든 개들의 집중을 방해하기 때문이죠. 마찬가지로 만약 반려견의 최고의 강화물이 수영인데 근처에 물이 없다면, 그것은 반려견의 최고의 강화물이 될 수 없습니다. 이런 환경에서 훈련할 때 아주 훌륭한 강화물입니다. 그래서 환경은 단지 방해 요소만이 아니라, 환경의 구조 자체가 무엇이며, 강화물로 활용할 수 있는 어떤 문제나 자산이 있는지가 중요합니다. 이제 고려해야 할 또 다른 점은, 당신이 훈련하려는 행동에 대해 가장 파격적이면서도 훌륭하고 수용 가능한 강화물이 무엇인가 하는 점입니다. 예를 들어, 정밀한 힐(heel) 위치를 훈련한다고 가정하면, 그 정밀한 힐 자세를 유지한 상태에서 개에게 바로 줄 수 있는 것을 선택할 수 있습니다. 수영하러 가는 것을 강화물로 사용할 수도 있겠지만, 그건 가끔씩밖에 할 수 없기 때문에 한계가 있을 것입니다. 반면에
Let's just talk about when we are training in a reinforcement-based program, meaning you've decided you want to train your dog without the use of physical corrections or verbal intimidation. You don't want to say, hey, no, ah. You don't want to give them a collar pop. When you decide to train that way, then your number one tool is reinforcement. And so, you've got to become brilliant at the use of and the understanding of reinforcement. What do you need to know? First of all, you have to know what are your dog's best reinforcements. That is number one, what is like off the charts, crazy, outrageous, oh my gosh, I'm losing my mind because this reinforcement is coming out. You also need to know what is a really good reinforcement. And then you've got to know what is an acceptable reinforcement. Now those things are going to change based on the age of the dog's experience and training and the location where you're training. So, what's outrageous, what's good, what's acceptable in this environment at this stage of your training. All right. And what about under these distractions? Because if there's a lot of environmental distractions, you're going to have to go even higher. So, what was previously acceptable won't be acceptable when there's a lot of distractions around. Every time you train your dog, you have to consider where you're training and what is outrageous, what is good, what is acceptable, because those are going to keep changing. You also have to know what are you limited by the environment you're training. So, for example, if you're training in a class, you can't use a squeaky toy because you're going to be just distracting all of the other dogs. Likewise, if your dog's outrageous number one reinforcement is swimming and there's no body of water nearby, that can't be your dog's number one outrageous outstanding reinforcement when you're training in this environment. So, the environment, it's not just the distractions, it's the structure of the environment, what problems or what assets does it present that you can use as a reinforcement. Now, the other thing you need to consider, what is the most outrageous, good and acceptable reinforcement for the behavior that you are training. So, if you were going to be training, let's say precision heel position, you may choose to use something that can be delivered to the mouth of the dog in that precision heel position. So, you could use like go for a swim to reinforce that, but it would be more limiting because it could only happen once every once in a
2:31
고가치의 간식을 주는 것은 훨씬 더 빠르게 반복할 수 있습니다. 그렇다면 그 행동은 무엇이며, 평소 상황에서 개가 가장 좋아하는 것을 사용하는 것이 합리적일까요? 좋습니다. 그것이 훈련하려는 행동에 효과가 있을까요? 다음 고려 사항은 그 강화물을 전달하는 방식입니다. 그리고 그것은 훈련 목표에 따라 크게 달라집니다. 당신의 훈련 목표는 무엇인가요? 만약 제가 개가 발톱 손질을 받아들이도록 모양을 만들어가고 싶다면, 고려할 점이 많지만 자세는 개가 옆으로 누운 상태가 될 것입니다. 저는 개에게 그 자세에서 벗어나도 좋다는 마커 단어를 사용하거나, 더 나아가 직접 강화물을 가져다 줄 것입니다. 즉, 강화물을 어디에 두느냐가 개가 누워 있는 행동을 실제로 강화하는 것입니다. 이 팟캐스트에는 많은 보석 같은 지식, 즉 '아하' 하는 순간이나 핵심적인 정보들이 담겨 있을 것입니다. 제 멘토인 밥 베일리는 항상 이렇게 말합니다. 강화는 과정이지, 찰나의 순간이 아니라고 말이죠. 그러니 '상관없어, 나는 마커를 사용하니까 개를 일으켜 간식을 주고 다시 눕게 하면 돼'라고 생각할 수 있습니다. 하지만 마커는 분리합니다. 그 순간 개가 하고 있는 반응입니다. 나중에 그 행동에 대해 동물이 받는 강화는 예를 들어 개를 일어나게 할 때, 마커와 개가 실제로 간식을 얻기까지 일어나는 모든 것을 강화합니다. 그 사이에 간격이 있죠? 그래서 강화 과정은 그 모든 것을 강화합니다. 큰 간격을 원하시나요, 아니면 신경 쓰지 않으시나요? 간격이 클수록 학습은 더 느려집니다. 클릭은 개의 행동을 끝내게 해서는 안 됩니다. 보셨나요? 밥 베일리의 사인입니다. 정말 멋지지 않나요? 이 점을 계속 주목해 주세요. 다음 2분 동안 저는 여러분의 반려견 훈련 코치가 되고 싶습니다. 반려견 훈련이 처음이시라면, 반려견 훈련이 완전히 처음이신 거죠. 분명 배울 점이 있을 겁니다. 그리고 만약 경험이 꽤 있으신 분이라도 스포츠 트레이너의 관점에서 반려견 훈련에 대한 새로운 통찰력을 얻으실 수 있을 거라고 확신합니다. 좋습니다. 자, 이 개가 앞발을 이 바위 위에 올리도록 훈련하고 싶어 하는 트레이너라고 가정해 봅시다. 그게 여러분의 과제입니다. 반려견 트레이너로서 준비되셨겠죠. 클릭기를 손에 쥐고 덕트 테이프로 거꾸로 감아 고정하세요. 간식 주머니는 아주 맛있는 간식으로 가득 채웠고요. 훈련 중에 머리카락이 눈을 가리지 않게 뒤로 묶었습니다. 시작하죠. 자, 바로 그 순간입니다. 개가 앞발을 바위에 올립니다. 트레이너는 클릭기를 눌러 개에게 그 순간을 표시합니다.
while where delivering high value food could happen more quickly. So, what is the behavior and is it reasonable to consider using what the dog loves most of all in a normal situation? All right. Does it work for the behavior that you are training? The next consideration is the delivery of that reinforcement. And that is really dependent upon your goal. What is your goal of training? So, if I was wanting to shape my dog to accept their nails being trimmed, there's a lot of things that come up to this, but the position would be they're lying on their side. And I would either give them a marker word, which meant you could get out out of that position or more likely what I would do is bring the reinforcement to them. So, the placement of the reinforcement actually reinforces them lying down because this is a gem and there's going to be a lot of aha moments or knowledge nuggets or gems in this podcast. My mentor Bob Bailey always says, reinforcement is a process. It's not a moment in time. So, you can say, well, it doesn't matter. I can get my dog up and feed them and then they can go back down because I'm using a marker. The marker isolates the response that the dog is doing at that moment. The reinforcement that the animal gets later for that behavior, like when you let them get up, reinforces everything that happens between the marker and them actually getting the cookie. There's a gap in there, right? So, the reinforcement process, it reinforces all of that. Do you want a large gap or do you care? The larger the gap, the slower the learning. The click should not end your dog's behavior. Did you see that? Autograph by Bob Bailey. How cool is that? Stick with me on this. For the next two minutes, I'd like to be your dog training coach. If you're brand new to dog training, you're brand new to dog training. I know you're going to learn something. And if you've been around the block a few times, I'm pretty confident you still might get some different insight from a sports trainer's viewpoint into your dog training. Okay. Let's just say you are a trainer who wants to train this dog to put these front paws onto this rock. That's your assignment. As a dog trainer, you're ready. You've got your clicker duct tape backwards to your hand. You've got your bait pouch filled with handy dandy tasty treats. Your hair's tied back so it won't get in your eyes when you're training. Let's go. Okay. The moment is upon us. The dog hits his front paws to the rock. The trainer clicks her clicker marking that moment for the
5:14
클릭 소리를 들은 개는 트레이너 앞으로 가서 위아래로 방방 뛰며 간식을 기다립니다. 심지어 간식 주머니에 무게를 실어 트레이너를 도우려 할지도 모릅니다. 트레이너는 개의 이런 노력을 눈치채지 못한 채 간식 주머니에 손을 넣어 간식을 꺼내고 잘한 일에 대해 보상합니다. 여러분은 "수잔, 뭐가 문제라는 거죠? 정말 환상적인 훈련 세션처럼 보였는데요"라고 하실지도 모릅니다. 글쎄요, 핵심은 이렇습니다. 강화는 하나의 과정입니다. 단지 시간 속의 고립된 순간이 아닙니다. 그럼요. 클릭커는 개의 앞발이 바위에 닿는 순간을 포착해 개에게 이렇게 말해준 거죠. 내가 찾는 게 바로 이거야. 이걸 더 많이 하면 보상을 받을 수 있어. 자, 이제 강화가 하나의 과정이라는 것을 알고, 그 과정을 살펴봅시다. 물론이죠. 클릭커는 개에게 네가 그 바위에 앞발을 올리는 게 마음에 든다고 말해준 거예요. 하지만 간식은 개가 궁극적으로 원했던 보상이고, 간식은 클릭 이후에 일어나는 모든 행동을 강화합니다. 왜냐하면 보상이 빠를수록, 동물이 얻는 행동에 대한 강화도 더 커지기 때문이죠. 이해가 가죠? 강화는 하나의 과정입니다. 마커를 사용하더라도 여전히 보상의 전달 위치를 고려해야 합니다. 예를 들어, 개가 당신으로부터 멀리 달려가도록 가르치고 싶어서 클릭을 했다면, 그건 좋습니다. 그런데 개가 다시 돌아와서 당신 앞에 앉았고, 당신은 주머니에서 간식을 꺼내고 있었다면요? 네, 당신은 개가 당신에게서 멀어지는 행동을 강화하는 것이기도 하지만, 동시에 돌아와서 당신 앞에 앉는 행동도 강화하게 되는 것인데, 이는 개가 당신에게서 멀어지게 하려는 의도와는 상반되는 것이죠. 그래서 당신에게서 멀리 달려가는 행동을 훈련하는 데 훨씬 더 오랜 시간이 걸릴 것입니다. 정말 중요하죠. 보상의 전달 위치는 매우 중요합니다. 왜냐하면 그 간격이 존재하기 때문이죠. 다음은 보상의 전달입니다. 음, 수잔, 보상의 전달과 위치는 무엇이 다른가요? 위치는 개가 무엇을 하고 있는가에 대한 것이고, 전달은 개에게 보상을 가져다줄 때 얼마나 신속하고 계획적으로 움직이는가에 대한 것입니다. 그래서 훈련 초기에는 개가 앉은 자세를 유지하도록 아주 빠르게 보상하고 싶을 수 있어요. 쾅, 쾅, 쾅, 쾅, 쾅, 쾅, 쾅 이렇게요. 하지만 강아지가 릴리스 신호를 들을 때까지 자세를 유지해야 한다는 것을 이해하게 되면, 저는 '쿠키'라고 말한 뒤 간식을 전달하기 전에 잠시 간식을 들어 보일 수도 있어요. 개가 앞발을 허우적거리기 시작하나요? 음, 만약 그렇다면 저는 '쿠키'라고 말하지 않을 거예요. 다음 차례를 위해서, 저는 그냥 간식을 높이 들고 제가 보상을 주기 전에 '네 선택이야(it's your choice)' 훈련을 좀 하려고 합니다. 좋아요. 그러니까 전달(delivery)과 배치(placement)는 다른 겁니다. 전달은 얼마나 빠르게
dog. Having heard the clicker, the dog then goes in front of the trainer, bouncing up and down, ready for his treat. Maybe even helping the trainer by digging his weight into the bait bag. The trainer, oblivious to the dog's effort, reaches into her bait bag, gets a treat and rewards the dog for a job well done. You may be saying, well, Susan, what's wrong with any of that? That looked like a fantastic training session. Well, here's the thing. Reinforcement is a process. It is not an isolated moment in time. So sure. The clicker isolated the moment the dog's paws hit the rock and told the dog, this is what I'm looking for. Do more of this and you can earn reinforcement. Now, knowing that reinforcement is a process, let's look at the process. Sure. The clicker told the dog, I like when you put your paws on that rock. However, the cookie is the reinforcement the dog ultimately wanted and the cookie reinforces all of those behaviors that happen after the click. Because the more rapid the reinforcement, the more reinforcement for a behavior that the animal's getting. Makes sense, right? Reinforcement is a process. If you use a marker, you still need to be considerate of the placement of the reinforcement. So, if you wanted to teach a dog to run far away from you and you marked, that's good. And then they came back and sat in front of you while you dug a cookie out of your pocket. Yeah, you would be reinforcing them leaving you, but you would also be reinforcing them for coming back and sitting in front of you, which would be in opposition to them running away from you. So, the behavior of running away from you would take a lot longer for you to train. Super important. The placement of reinforcement is critical because you've got that gap. Next is the delivery of reinforcement. Well, what's the difference, Susan? The delivery and the placement? The placement is what the dog is doing. The delivery is how swift and deliberate you are when you are bringing that reinforcement to them. So, early on in training, I may want to reinforce a dog for holding a sit like super fast, right? Boom, boom, boom, boom, boom, boom, boom. But as the puppy understands that the idea is hold position until you hear a release word, I may say cookie and hold up the cookie just for a moment before I deliver it to them. Do they start paddling their paws? Well, if they do, I'm not going to say cookie for the next one. I'm just going to hold it up and do some, it's your choice before I deliver the cookie. All right. So, the delivery and the placement are different. The delivery is how quickly you're
7:58
개에게 보상을 주는지, 강화 과정의 간극을 얼마나 길고 크게 만드는지에 대한 것이죠. 그리고 배치는 개가 수행하고 있는 최종 행동에 얼마나 많은 가치를 부여하느냐입니다. 알겠죠? 강화 과정, 이건 하나의 과정입니다. 그래서 사람은 가치, 전달, 그리고 배치에 신경을 써야 하고, 개는 어느 정도 편하게 훈련하는 셈이죠.
getting it to the dog, how long and big you grow that reinforcement process gap. And the placement is how much value are you bringing to the final, the behavior that the dog is doing. Okay. Reinforcement process. It's a process. So, the human's got to be concerned with the value, the delivery, and the placement, the dog, they've got it kind of easy.